Model performance
Beat the closing line
+0.58pp
813 picks published as value, settled, at a real recorded price.95% interval +0.03 to +1.01pp over 337 picks across 29 match-days. The interval excludes zero.
It clears zero by 0.03pp. The smallest effect this sample could reliably detect is 0.70pp (80% power, 337 picks across 29 match-days). The measured 0.58pp is below that, so this establishes that we are not BEHIND the close; it does not establish how far ahead.
Win rate
55.8%
737 decided picks of 813 settled. A win rate is not a return: the price each pick was taken at decides that.
Return per pick
-1.36%
At the price each pick was available at. 95% interval -9.71 to +6.89% — wide enough that no return can be claimed from it either way.
Shopped to the best price we observed on the ladder, the same picks return +5.39u (+0.66%) — a change in the price, not in the results. 412 of 813 picks had a priced ladder at their own line; the rest keep the price we recorded. We name the book on each pick, because which books you can actually use is not something we can know for you.
85% of these picks were published before the current settings. We changed what the model is allowed to publish on 15 September 2026, so the record above mostly measures picks the gate would select differently today. It is a true record of what we called; it is not yet a record of what we are calling now. Enough picks from the current settings will have settled to report them separately around 25 September, and to compare them properly around 14 October.
These two point opposite ways, and both are true. Beating the closing line says the forecast is better than the market’s. The return is what is left after the bookmaker’s margin is paid out of that edge — and the margin is currently several times larger than the edge, so a better-than-market forecast still loses money at the prices available.
What this counts
Published as a value pick, settled, at a real recorded entry price, graded forward rather than by a retrospective sweep.
Sold as value today: asian handicap — +0.95pp (+0.11 to +1.54pp) over 387 picks.
A previously published figure read 56.3% over 583 picks. Same picks, different population. The headline keeps markets that were later retired, and counts only picks carrying a real recorded price. No result changed.
What this counts
The record counts markets currently published as value picks. Markets we no longer publish as value picks, and in-play calls, are still graded and still shown on their own fixtures — they just do not move this number.
Those 3,610 retired picks returned -6.3%. We retired those markets on this same history, so read the scoped figure as a description of what we publish now, not as an independent forecast.
A correction, stated in full. A defect let 120 picks that had failed our own edge floor be counted as value picks. They returned -26.4%, so removing them IMPROVES the figure above: as published it was -1.4% over 707 picks. We publish all three so the correction can be checked rather than trusted. The rule was written down before it was applied, it judges each pick against the floor that was in force when that pick was published, and it leaves in 22 settled picks we could not demonstrate the defect for — which were 19-3 in our favour.
Every pick ever settled, all markets: -4.7% (-174.8u over 3,702 priced picks).
"Live forward" picks were graded as their matches finished; "backfilled" picks were settled after the fact by a bulk sweep. Graded between 2026-06-16 and 2026-09-27.
Methodology
How we keep this honest
• Only picks we actually published to you are counted — the exact picks you can act on, nothing cherry-picked.
• Pre-match picks lock at kickoff and are graded win or lose — we never quietly delete a losing pick.
• Win rate is shown against break-even — the win rate the odds require to profit — so the number means something on its own.
• "Beat the closing line" is how often we secured a better price than where the market ended up — the sharpest sign of a real edge.
• Return is measured at fair (no-margin) odds, isolating our forecasting skill from the bookmaker's cut.
• Pushes return the stake — a line landing exactly counts as neither win nor loss — and a cancelled match voids its picks instead of leaving them hanging.
• Small samples are flagged, not hidden, and every graded pick is listed in the archive below.
Calibration curve
Predicted vs actual
When we say a chance is X%, it should happen about X% of the time. On average our stated chances land within ±3.4 points of what actually happens.
Measured on published picks only, which are selected for value — so this is a stricter test than a model's raw output.
Forward calibration
Do we deliver what we claim?
Measured on picks we published and settled promptly — not on the data the model was fitted to. A negative figure means the model claimed more than it delivered.
Across every published pick the model claims more than it delivers, and that is established. Whether the picks we sell as value are over-confident is not: that sample is smaller and its interval still contains zero.
Bins: equal width deciles · all history. The error figure moves with the bin choice, so the choice is stated. Intervals are a clustered bootstrap over match-days, because picks on one slate share fixtures and model version.
These are three nested samples, widest first — not a trend. Reading a direction across them is how a difference of half a point once got published as a finding.
Closing-line value
Did the price move toward us?
A pick beats the close when the market shortens it after we publish. A price that never moved counts as half, not as a loss, and a pick whose close we captured before publishing is left out entirely — it could not have been beaten. It is a faster read on whether a market is worth selling than the settled return, because it uses every pick rather than only the decided money.
14 outcomes · 8 match-days · 95% interval 42.9–86.7% · too few to conclude
188 outcomes · 22 match-days · 95% interval 50.9–63.9%
20 outcomes · 7 match-days · too few to conclude
8 outcomes · 3 match-days · too few to conclude
57 outcomes · 13 match-days · 95% interval 37.8–53.8% · too few to conclude
36 outcomes · 7 match-days · too few to conclude
Measured on picks published as value, settled forward against a captured closing price, pre-match only. Above 50% the market moved toward the pick after we published it.
And is it moving?
All markets pooled, newest last. A calendar week holds too few match-days at our volume to carry a rate, so each point covers a rolling window and the series steps weekly.
331 outcomes · 28 match-days · 95% interval 48.0–59.7%
331 outcomes · 28 match-days · 95% interval 48.0–59.7%
331 outcomes · 28 match-days · 95% interval 48.0–59.7%
330 outcomes · 27 match-days · 95% interval 47.8–59.5%
327 outcomes · 26 match-days · 95% interval 47.1–58.9%
325 outcomes · 25 match-days · 95% interval 47.1–58.8%
325 outcomes · 25 match-days · 95% interval 47.1–58.8%
323 outcomes · 24 match-days · 95% interval 47.5–59.0%
90 days ending on each point's own date, so consecutive points overlap and a single match-day cannot move the line by itself. Stepped weekly because a calendar week at our current volume holds too few match-days to carry a rate at all.
By confidence band
Hit rate by confidence
Realistic range 34–100%
Realistic range 54–64%
Realistic range 44–58%
Reliability by slice
How much the model beats the base rate
Measured nightly over every graded outcome, not only the picks we published. Each figure is how much sharper our probabilities are than simply knowing how often the thing happens; zero means no sharper at all. Every row states the sample it was measured over and its interval, and a slice with too few outcomes is shown and labelled rather than quietly dropped.
By market
11 slicesEvery graded outcome, not only the picks we published.
9,081 outcomes · 18 match-days · 95% interval 87.3–89.8%
claimed 97.4% · happened 88.6%
12,769 outcomes · 64 match-days · 95% interval 12.5–13.7%
claimed 13.5% · happened 13.1%
3,408 outcomes · 65 match-days · 95% interval 48.2–51.6%
Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.
40,131 outcomes · 44 match-days · 95% interval 49.5–50.5%
Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.
5,144 outcomes · 65 match-days · 95% interval 32.1–34.7%
claimed 33.9% · happened 33.4%
3,420 outcomes · 64 match-days · 95% interval 41.8–47.0%
Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.
5,140 outcomes · 64 match-days · 95% interval 65.4–67.9%
claimed 66.5% · happened 66.7%
33,181 outcomes · 65 match-days · 95% interval 49.4–50.4%
Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.
By tier
4 slicesHow far into a fixture the read was taken.
9,081 outcomes · 18 match-days · 95% interval 87.3–89.8%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
46,359 outcomes · 58 match-days · 95% interval 44.7–46.1%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
81,572 outcomes · 62 match-days · 95% interval 44.8–46.0%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
24,400 outcomes · 64 match-days · 95% interval 39.2–40.4%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
By competition
28 slicesCompetitions with enough settled outcomes to measure.
387 outcomes · 4 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
432 outcomes · 3 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
375 outcomes · 3 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
414 outcomes · 2 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
432 outcomes · 2 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
624 outcomes · 2 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
129 outcomes · 1 match-days
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
12,472 outcomes · 17 match-days · 95% interval 45.5–47.6%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
By confidence band
3 slicesThe band is a threshold on a score that does not currently rank outcome.
3,377 outcomes · 64 match-days · 95% interval 46.9–52.5%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
5,533 outcomes · 53 match-days · 95% interval 47.7–52.6%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
433 outcomes · 28 match-days · 95% interval 35.0–50.8%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
By price
5 slicesThe price the pick was available at, not the fair price.
1,040 outcomes · 59 match-days · 95% interval 33.9–40.7%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
1,979 outcomes · 61 match-days · 95% interval 46.0–51.0%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
1,845 outcomes · 49 match-days · 95% interval 12.1–16.5%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
3,219 outcomes · 61 match-days · 95% interval 66.2–72.2%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
1,260 outcomes · 41 match-days · 95% interval 55.7–70.6%
This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.
By market and band
30 slicesThe finest cut, and the one where cells go thin.
102 outcomes · 23 match-days · 95% interval 43.3–64.2%
170 outcomes · 18 match-days · 95% interval 2.9–18.6%
215 outcomes · 2 match-days
1,010 outcomes · 35 match-days · 95% interval 49.2–57.1%
1,249 outcomes · 40 match-days · 95% interval 44.9–51.7%
878 outcomes · 49 match-days · 95% interval 45.7–52.3%
120 outcomes · 18 match-days · 95% interval 31.7–54.1%
1,956 outcomes · 51 match-days · 95% interval 48.3–56.6%
By sport
By market
By analysis tier
Hit rate by tier
By edge bucket
Hit rate by edge range
By league
Hit rate by competition
By model version
Performance by model
Model audit log
Version history
matchsense-basketball-spread v2026.03.1
Model matchsense-basketball-spread v2026.03.1 is currently active
9/29/2026
matchsense-basketball-spread v2026.03.1
Model matchsense-basketball-spread v2026.03.1 last trained
9/29/2026
matchsense-football-moneyline v2026.03.1
Model matchsense-football-moneyline v2026.03.1 is currently active
9/29/2026
matchsense-football-moneyline v2026.03.1
Model matchsense-football-moneyline v2026.03.1 last trained
9/29/2026
matchsense-basketball-spread v2026.03.1
Model matchsense-basketball-spread v2026.03.1 created for BASKETBALL Spread
3/25/2026
matchsense-football-moneyline v2026.03.1
Model matchsense-football-moneyline v2026.03.1 created for FOOTBALL Moneyline
3/25/2026