Skip to main content
News
FOOTBALLmethodology07/08/20264 min readMatchSense Analytics Team

How We Grade and Report Prediction Accuracy

Accuracy claims are only meaningful if the measurement method is public. This page documents exactly how MatchSense grades predictions, which metrics we track, how the sample is defined, and which commonly advertised numbers we refuse to publish because they are misleading.

Key takeaways

  • Every published prediction is graded against the final result, including the ones that lose.
  • Brier score and log loss measure probability quality directly; win rate does not.
  • Calibration curves are reported by confidence band, so overconfidence is visible rather than hidden in an average.
  • The figure available at kick-off is tracked as external evidence that the model's assessment holds an edge independently of results.
  • We do not publish win rate without the accompanying probability, cherry-picked streaks, or ROI computed at figures that were not available at publication.

Any service can claim accuracy. Very few publish the method by which they measure it, which is what makes the claim checkable. This page is that method.

What gets graded

Every prediction published on MatchSense is recorded before kick-off with its market, selection, probability, the associated number available at publication, and a timestamp. After the match, it is graded against the final result.

Nothing is excluded. Losing predictions are graded identically to winning ones, and no competition, market or period is dropped from the sample after the fact. The reason is simple: the moment a record becomes selective, every metric computed from it becomes meaningless, and that is the single most common way accuracy claims are inflated.

The metrics we track

Brier score. The mean squared difference between the predicted probability and the outcome, where the outcome is 1 if the event happened and 0 if it did not. Lower is better; a perfect forecaster scores 0. Its important property is that it punishes confident errors far more than uncertain ones — predicting 95% on something that does not happen costs much more than predicting 55%.

Log loss. A stricter alternative that penalises confident errors even more aggressively. It is useful precisely because it is unforgiving of the failure mode that matters most: overconfidence.

Calibration by band. Predictions are grouped into confidence bands and the realised hit rate in each band is reported alongside the predicted rate:

Predicted bandPredicted rateRealised rate
50-60%55%reported per period
60-70%65%reported per period
70-80%75%reported per period
80%+85%reported per period

Reporting by band matters because an aggregate average can hide systematic overconfidence at the high end, which is exactly where the model's most confident predictions are made and exactly where a miscalibrated model does the most damage.

What we deliberately do not publish

Win rate without the accompanying probability. A hit rate quoted alone is not an accuracy measure. It can be trivially inflated by predicting heavy favourites, and it says nothing about whether following the predictions is profitable.

Streaks. "Nine winners in a row" describes variance. Every strategy with a positive win rate produces streaks, and so does every strategy with a negative one.

ROI at unavailable figures. Return calculated at the best figure found anywhere after the fact, rather than at the figure available when the prediction was published, is a fiction. Ours is computed at recorded publication figures.

Aggregate accuracy across markets. Pooling a 1X2 hit rate with a total goals hit rate produces a number that means nothing, because the base rates differ. Metrics are reported per market.

What we would expect you to demand

The checklist in how to evaluate a football prediction service was written to be applied to us as much as to anyone else. A complete timestamped record, a sample of several hundred predictions, ROI at the recorded figures, and calibration by band. If any of those is missing from a service you are paying for, ask why.

Why this is published at all

There is a commercial argument against publishing a method that makes your own numbers falsifiable. We take the opposite view: a prediction service whose accuracy cannot be checked is indistinguishable from one that is guessing, and the only durable way to be distinguishable is to be checkable.

The modelling pipeline that produces the predictions being graded here is documented in how the MatchSense prediction model works.

Related reading

Frequently asked questions

How is a prediction graded?

Against the final result of the match for the market it was published in, with the timestamp and the figure available at publication recorded alongside it. Predictions are graded whether they win or lose, and the graded record is what feeds the accuracy metrics.

What is a Brier score?

Brier score is the mean squared difference between predicted probabilities and actual outcomes, where the outcome is 1 if the event occurred and 0 if it did not. Lower is better, a perfect forecaster scores 0, and it penalises confident wrong predictions much more heavily than uncertain ones.

Why not just publish a win percentage?

Because a win percentage without the accompanying figure is uninterpretable. A 70 percent hit rate on heavy favourites loses money, while a 30 percent hit rate at low probabilities can be highly profitable. Win rate quoted alone is a marketing number rather than an accuracy measure.

What sample do the accuracy metrics cover?

All published predictions in the stated period and market, with no exclusions. Removing losing predictions, or reporting only a favourable subset of competitions, invalidates the metric entirely.

Get the weekly briefing

Next week's fixtures, how last week's published picks actually settled, and the newest analysis. One email a week. Unsubscribe in one click.

We never share your address, and every email has a one-click unsubscribe.

Related reading

Play responsibly. MatchSense provides sports analysis for informational purposes only and does not accept wagers. Need support? Call 1-800-522-4700.

How MatchSense Grades and Reports Prediction Accuracy