Skip to main content
Methodology
Methodology

How MatchSense builds predictions

Every number shown in MatchSense comes from a structured pipeline. No black boxes, no vague claims, just data, models, and transparent outputs.

Model Pipeline

XGBoost + LightGBM ensemble, priced from the scoreline distribution

footballbasketballExact derivation from the scoreline distribution (no sampling error)

Data Sources

Live odds feeds from configured providers

Historical match results and team statistics

Confirmed lineups and availability, where the team sheet has been published and captured (77% of fixtures pre-kickoff)

In-play pace, possession, and pressure metrics

Refresh Cadence

Predictions are refreshed at each analysis tier as new data becomes available.

Prematch Early

Generated when match enters catalog

Prematch Lineup

Generated once a confirmed team sheet has been published and captured; fixtures without one stay on the early-tier read

In-Play

Refreshed on live match state changes

Calibration

Probabilities are calibrated against historical outcomes, and calibration is measured and published PER MARKET rather than assumed uniform. It is not uniform: selecting picks for publication maximises estimated edge, and estimated edge is largest where a probability is most overconfident, so the published population is measurably less well calibrated than the model's full output.

Trust Center

MatchSense is built on verifiable outputs. Every claim on this platform can be traced.

Confidence bands: withdrawn

We used to print a HIGH / MEDIUM / LOW band beside every pick. Measured on 19 September 2026 over 416 settled picks across 27 match-days, it did not rank anything: LOW returned +7.7 units from 306 picks and MEDIUM lost 5.2 from 110, and neither the difference in return nor the difference in closing-line value excluded zero once the sample was resampled by match-day. We removed it rather than keep a label a reader would reasonably act on. It is not being refined quietly either — the ensemble-agreement figures it would be rebuilt from are recorded on none of the settled picks in the two markets we sell.

Edge Interpretation

Edge is the gap between the model's estimated probability and the market-implied probability, and it is what selects which picks we publish. Read it as an estimate rather than a forecast: selecting on an estimate selects its error too, so published picks have historically run about four percentage points over-confident. That is why every edge and EV we show carries the return those picks have actually made.

Projected Scores

Projected scores are the expected value of the scoreline distribution, not a single outcome prediction. The distribution is derived exactly from the model's goal rates rather than sampled from them, so the figure carries no simulation error.

Tier Refreshes

When new data arrives, predictions are regenerated. The archive shows exactly what changed between each tier so you can trace the model's reasoning evolution.

MatchSense does not guarantee outcomes. All analysis represents the model's best estimate given available data.