Skip to main content
Performance
Performance

Model performance

Beat the closing line

+0.58pp

813 picks published as value, settled, at a real recorded price.95% interval +0.03 to +1.01pp over 337 picks across 29 match-days. The interval excludes zero.

It clears zero by 0.03pp. The smallest effect this sample could reliably detect is 0.70pp (80% power, 337 picks across 29 match-days). The measured 0.58pp is below that, so this establishes that we are not BEHIND the close; it does not establish how far ahead.

Win rate

55.8%

737 decided picks of 813 settled. A win rate is not a return: the price each pick was taken at decides that.

Return per pick

-1.36%

At the price each pick was available at. 95% interval -9.71 to +6.89% — wide enough that no return can be claimed from it either way.

Shopped to the best price we observed on the ladder, the same picks return +5.39u (+0.66%) — a change in the price, not in the results. 412 of 813 picks had a priced ladder at their own line; the rest keep the price we recorded. We name the book on each pick, because which books you can actually use is not something we can know for you.

85% of these picks were published before the current settings. We changed what the model is allowed to publish on 15 September 2026, so the record above mostly measures picks the gate would select differently today. It is a true record of what we called; it is not yet a record of what we are calling now. Enough picks from the current settings will have settled to report them separately around 25 September, and to compare them properly around 14 October.

These two point opposite ways, and both are true. Beating the closing line says the forecast is better than the market’s. The return is what is left after the bookmaker’s margin is paid out of that edge — and the margin is currently several times larger than the edge, so a better-than-market forecast still loses money at the prices available.

What this counts

Published as a value pick, settled, at a real recorded entry price, graded forward rather than by a retrospective sweep.

Sold as value today: asian handicap — +0.95pp (+0.11 to +1.54pp) over 387 picks.

A previously published figure read 56.3% over 583 picks. Same picks, different population. The headline keeps markets that were later retired, and counts only picks carrying a real recorded price. No result changed.

What this counts

The record counts markets currently published as value picks. Markets we no longer publish as value picks, and in-play calls, are still graded and still shown on their own fixtures — they just do not move this number.

Those 3,610 retired picks returned -6.3%. We retired those markets on this same history, so read the scoped figure as a description of what we publish now, not as an independent forecast.

A correction, stated in full. A defect let 120 picks that had failed our own edge floor be counted as value picks. They returned -26.4%, so removing them IMPROVES the figure above: as published it was -1.4% over 707 picks. We publish all three so the correction can be checked rather than trusted. The rule was written down before it was applied, it judges each pick against the floor that was in force when that pick was published, and it leaves in 22 settled picks we could not demonstrate the defect for — which were 19-3 in our favour.

Every pick ever settled, all markets: -4.7% (-174.8u over 3,702 priced picks).

Record provenanceLive forward 472 picks · +7.5% returnBackfilled 111 picks · +29.3% return

"Live forward" picks were graded as their matches finished; "backfilled" picks were settled after the fact by a bulk sweep. Graded between 2026-06-16 and 2026-09-27.

Methodology

How we keep this honest

• Only picks we actually published to you are counted — the exact picks you can act on, nothing cherry-picked.

• Pre-match picks lock at kickoff and are graded win or lose — we never quietly delete a losing pick.

• Win rate is shown against break-even — the win rate the odds require to profit — so the number means something on its own.

• "Beat the closing line" is how often we secured a better price than where the market ended up — the sharpest sign of a real edge.

• Return is measured at fair (no-margin) odds, isolating our forecasting skill from the bookmaker's cut.

• Pushes return the stake — a line landing exactly counts as neither win nor loss — and a cancelled match voids its picks instead of leaving them hanging.

• Small samples are flagged, not hidden, and every graded pick is listed in the archive below.

Calibration curve

Predicted vs actual

When we say a chance is X%, it should happen about X% of the time. On average our stated chances land within ±3.4 points of what actually happens.

0%-10%1 predictions
0%
10%-20%1 predictions
0%
20%-30%2 predictions
50%
30%-40%22 predictions
32%
40%-50%61 predictions
49%
50%-60%244 predictions
57%
60%-70%179 predictions
62%
70%-80%27 predictions
59%
80%-90%10 predictions
50%
90%-100%2 predictions
100%

Measured on published picks only, which are selected for value — so this is a stricter test than a model's raw output.

Forward calibration

Do we deliver what we claim?

Measured on picks we published and settled promptly — not on the data the model was fitted to. A negative figure means the model claimed more than it delivered.

Every published pick-3.79pp
2,934 picks · 41 match-days95% interval -6.48pp to -1.49ppECE 3.85pp
Picks sold as value-2.85pp
754 picks · 31 match-days95% interval -8.37pp to +2.49pp — contains zeroECE 4.01pp
Asian handicap, sold as value-2.21pp
356 picks · 29 match-days95% interval -11.06pp to +5.31pp — contains zeroECE 2.60pp

Across every published pick the model claims more than it delivers, and that is established. Whether the picks we sell as value are over-confident is not: that sample is smaller and its interval still contains zero.

Bins: equal width deciles · all history. The error figure moves with the bin choice, so the choice is stated. Intervals are a clustered bootstrap over match-days, because picks on one slate share fixtures and model version.

These are three nested samples, widest first — not a trend. Reading a direction across them is how a difference of half a point once got published as a finding.

Closing-line value

Did the price move toward us?

A pick beats the close when the market shortens it after we publish. A price that never moved counts as half, not as a loss, and a pick whose close we captured before publishing is left out entirely — it could not have been beaten. It is a faster read on whether a market is worth selling than the settled return, because it uses every pick rather than only the decided money.

Draw no bet64.3%

14 outcomes · 8 match-days · 95% interval 42.9–86.7% · too few to conclude

Asian handicap58.5%

188 outcomes · 22 match-days · 95% interval 50.9–63.9%

Moneyline 1x257.5%

20 outcomes · 7 match-days · too few to conclude

Double chance50.0%

8 outcomes · 3 match-days · too few to conclude

Over under corners46.5%

57 outcomes · 13 match-days · 95% interval 37.8–53.8% · too few to conclude

Over under goals38.9%

36 outcomes · 7 match-days · too few to conclude

Measured on picks published as value, settled forward against a captured closing price, pre-match only. Above 50% the market moved toward the pick after we published it.

And is it moving?

All markets pooled, newest last. A calendar week holds too few match-days at our volume to carry a rate, so each point covers a rolling window and the series steps weekly.

2026-09-2154.8%

331 outcomes · 28 match-days · 95% interval 48.0–59.7%

2026-09-2254.8%

331 outcomes · 28 match-days · 95% interval 48.0–59.7%

2026-09-2354.8%

331 outcomes · 28 match-days · 95% interval 48.0–59.7%

2026-09-2454.7%

330 outcomes · 27 match-days · 95% interval 47.8–59.5%

2026-09-2554.4%

327 outcomes · 26 match-days · 95% interval 47.1–58.9%

2026-09-2654.2%

325 outcomes · 25 match-days · 95% interval 47.1–58.8%

2026-09-2754.2%

325 outcomes · 25 match-days · 95% interval 47.1–58.8%

2026-09-2854.2%

323 outcomes · 24 match-days · 95% interval 47.5–59.0%

90 days ending on each point's own date, so consecutive points overlap and a single match-day cannot move the line by itself. Stepped weekly because a calendar week at our current volume holds too few match-days to carry a rate at all.

By confidence band

Hit rate by confidence

HIGHtoo few to trust2/2 (100.0%)

Realistic range 34–100%

LOW202/342 (59.1%)

Realistic range 54–64%

MEDIUM105/205 (51.2%)

Realistic range 44–58%

Reliability by slice

How much the model beats the base rate

Measured nightly over every graded outcome, not only the picks we published. Each figure is how much sharper our probabilities are than simply knowing how often the thing happens; zero means no sharper at all. Every row states the sample it was measured over and its interval, and a slice with too few outcomes is shown and labelled rather than quietly dropped.

By market

11 slices

Every graded outcome, not only the picks we published.

Player prop-3.3%−8.7pp off

9,081 outcomes · 18 match-days · 95% interval 87.3–89.8%

claimed 97.4% · happened 88.6%

Half time full time6.7%−0.4pp off

12,769 outcomes · 64 match-days · 95% interval 12.5–13.7%

claimed 13.5% · happened 13.1%

Both teams to score22.1%

3,408 outcomes · 65 match-days · 95% interval 48.2–51.6%

Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.

Over under corners22.7%

40,131 outcomes · 44 match-days · 95% interval 49.5–50.5%

Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.

Moneyline 1x229.4%−0.5pp off

5,144 outcomes · 65 match-days · 95% interval 32.1–34.7%

claimed 33.9% · happened 33.4%

Team totals29.5%

3,420 outcomes · 64 match-days · 95% interval 41.8–47.0%

Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.

Double chance29.9%+0.2pp off

5,140 outcomes · 64 match-days · 95% interval 65.4–67.9%

claimed 66.5% · happened 66.7%

Over under goals31.2%

33,181 outcomes · 65 match-days · 95% interval 49.4–50.4%

Both sides of this market are counted together, so over- and under-confidence cancel and no direction can be read from it.

By tier

4 slices

How far into a fixture the read was taken.

(none)-3.3%

9,081 outcomes · 18 match-days · 95% interval 87.3–89.8%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

PREMATCH LINEUP29.9%

46,359 outcomes · 58 match-days · 95% interval 44.7–46.1%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

PREMATCH EARLY30.2%

81,572 outcomes · 62 match-days · 95% interval 44.8–46.0%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

INPLAY50.8%

24,400 outcomes · 64 match-days · 95% interval 39.2–40.4%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

By competition

28 slices

Competitions with enough settled outcomes to measure.

AF536-11.3%

387 outcomes · 4 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

AF262-8.7%

432 outcomes · 3 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

AF5-7.9%

375 outcomes · 3 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

AF44-5.0%

414 outcomes · 2 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

AF89-4.2%

432 outcomes · 2 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

AF141-3.7%

624 outcomes · 2 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

AF254-3.0%

129 outcomes · 1 match-days

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

DED27.8%

12,472 outcomes · 17 match-days · 95% interval 45.5–47.6%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

By confidence band

3 slices

The band is a threshold on a score that does not currently rank outcome.

LOW24.4%

3,377 outcomes · 64 match-days · 95% interval 46.9–52.5%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

MEDIUM25.6%

5,533 outcomes · 53 match-days · 95% interval 47.7–52.6%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

HIGH39.5%

433 outcomes · 28 match-days · 95% interval 35.0–50.8%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

By price

5 slices

The price the pick was available at, not the fair price.

2.2-3.0-1.3%

1,040 outcomes · 59 match-days · 95% interval 33.9–40.7%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

1.8-2.2-0.1%

1,979 outcomes · 61 match-days · 95% interval 46.0–51.0%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

3.0+1.7%

1,845 outcomes · 49 match-days · 95% interval 12.1–16.5%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

<1.88.6%

3,219 outcomes · 61 match-days · 95% interval 66.2–72.2%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

(unpriced)50.9%

1,260 outcomes · 41 match-days · 95% interval 55.7–70.6%

This counts every market at once, and markets with two opposite sides cancel each other out, so no direction can be read from it.

By market and band

30 slices

The finest cut, and the one where cells go thin.

Team totals · LOW0.7%

102 outcomes · 23 match-days · 95% interval 43.3–64.2%

Correct score · MEDIUM1.6%−1.0pp off

170 outcomes · 18 match-days · 95% interval 2.9–18.6%

Half time full time · MEDIUM8.2%+1.7pp off

215 outcomes · 2 match-days

Over under corners · MEDIUM13.7%

1,010 outcomes · 35 match-days · 95% interval 49.2–57.1%

Over under goals · MEDIUM16.6%

1,249 outcomes · 40 match-days · 95% interval 44.9–51.7%

Over under goals · LOW18.2%

878 outcomes · 49 match-days · 95% interval 45.7–52.3%

Both teams to score · MEDIUM20.0%

120 outcomes · 18 match-days · 95% interval 31.7–54.1%

Asian handicap · MEDIUM21.9%

1,956 outcomes · 51 match-days · 95% interval 48.3–56.6%

By sport

BASKETBALL
25.0% hit4 graded
FOOTBALL
56.5% hit545 graded

By market

asian handicap
55.7% hit433 graded
over under corners
58.6% hit116 graded

By analysis tier

Hit rate by tier

PREMATCH EARLY250/442 (56.6%)
PREMATCH LINEUP59/107 (55.1%)

By edge bucket

Hit rate by edge range

0-3%9/14 (64.3%)
3-6%141/269 (52.4%)
6-10%137/237 (57.8%)
10%+22/28 (78.6%)

By league

Hit rate by competition

WC
61.2% hit103 graded
MLS
61.8% hit68 graded
ELC
50.7% hit67 graded
PD
57.1% hit56 graded
EL
70.0% hit40 graded
FL1
43.6% hit39 graded
PPL
50.0% hit34 graded
DED
57.6% hit33 graded
SA
58.1% hit31 graded
PL
36.7% hit30 graded
BL1
53.8% hit26 graded
CL
75.0% hit12 graded
ID1
80.0% hit5 graded
NBA
25.0% hit4 graded
AF239
0.0% hit1 graded

By model version

Performance by model

matchsense-football-moneyline v2026.03.1308/545 (56.5%)
matchsense-basketball-spread v2026.03.11/4 (25.0%)

Model audit log

Version history

matchsense-basketball-spread v2026.03.1

Model matchsense-basketball-spread v2026.03.1 is currently active

activated

9/29/2026

matchsense-basketball-spread v2026.03.1

Model matchsense-basketball-spread v2026.03.1 last trained

trained

9/29/2026

matchsense-football-moneyline v2026.03.1

Model matchsense-football-moneyline v2026.03.1 is currently active

activated

9/29/2026

matchsense-football-moneyline v2026.03.1

Model matchsense-football-moneyline v2026.03.1 last trained

trained

9/29/2026

matchsense-basketball-spread v2026.03.1

Model matchsense-basketball-spread v2026.03.1 created for BASKETBALL Spread

created

3/25/2026

matchsense-football-moneyline v2026.03.1

Model matchsense-football-moneyline v2026.03.1 created for FOOTBALL Moneyline

created

3/25/2026