- Bankroll
- $1000.00
- Today
- +0.00
- Risk
- 0.0%
- Open
- 0
- 7-day
- +0.00
Model vs the closing line
Every other number on this site scores the model against itself — against the outcomes it predicted, or against a fair price derived from its own probability. This page scores it against someone else's opinion: the closing market on the same matches. It is the only measurement here that predicts whether betting the model makes money, and the only one whose answer can be no.
Walk-forward floor (Elo only)
3,500 matches · 2024-08-15 → 2026-07-17Market beats the model
Over 3500 matches the model's Brier is 0.0271 WORSE than the closing line's (± 0.0025). This is the expected result: the close is a strong forecast. It means flat betting this model at market prices loses.
- Model
- 0.6021
- Closing line
- 0.5750
- Difference
- +0.0271market better
- Model
- 1.0083
- Closing line
- 0.9670
- Difference
- +0.0414market better
- Model
- 50.9%
- Closing line
- 54.4%
- Difference
- −3.5ppmarket better
Is the difference real?
+0.0271 ± 0.0025 (10.8σ)
Beyond 2 standard errors, so the gap is not sampling noise at this sample size.
Model closer to the truth
36.3%
Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.
Elo-only with ratings frozen as of each match — no calibration, no news, no market. The lowest defensible version of the model, so a loss here sizes the gap the live signal stack would have to close rather than condemning it.
Live model (Platt + news)
187 matches · 2026-04-26 → 2026-07-23Not enough matches to say
187 matches with both a model probability and a closing line — below the 200 needed before a comparison means anything. No verdict.
- Model
- 0.6218
- Closing line
- 0.6024
- Difference
- +0.0194market better
- Model
- 1.0430
- Closing line
- 1.0061
- Difference
- +0.0369market better
- Model
- 50.3%
- Closing line
- 50.8%
- Difference
- −0.5ppmarket better
Is the difference real?
+0.0194 ± 0.0157 (1.2σ)
Below the 200-match floor for publishing a verdict, so no claim either way.
Model closer to the truth
48.7%
Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.
The probabilities the site actually published: Elo plus Platt calibration and news features. Market-blind — the market snapshot only ever enters the edge calculation, never the probability.
Where these prices came from
- Closing lines held
- 3,550
- Matches compared
- 3,513
- Excluded
- 37
- Mean overround
- 3.87%
Price source
- fd-co-uk:pinnacle-close2,631 matches
- fd-co-uk:market-avg-close882 matches
Why 37 were excluded
- No model prediction
- 37
- Analysis after kickoff
- 1
Reading this honestly
- Lower is better for both Brier and log loss — they measure how wrong a forecast was. The side with the smaller number is the better forecaster.
- The market is expected to win. A devigged closing price from a sharp book aggregates far more information than an Elo model. Beating it is rare, and a claim to have done so deserves more scepticism than a claim not to have.
- The verdict is decided on Brier, not log loss. Both are proper scores, but log loss is unbounded — one confident miss can move its mean more than the rest of the sample combined. Brier is bounded, so its paired standard error is a trustworthy noise floor. Both are shown so you can disagree.
- The market is devigged multiplicatively, which slightly understates the favourite. Sharpest-book columns are preferred to keep that bias small.
- A season can mix price sources. football-data.co.uk stopped filling Pinnacle closing odds in January 2026, so later matches fall back to the market average — measured overround 6.2% against Pinnacle's 3.0%. A wider-margin book is a slightly weaker benchmark, and the bookmaker list shows the split per season.
- Both prediction sources are market-blind: walk-forward is Elo-only with frozen ratings, and the live model's probabilities never read the market snapshot.
- A verdict of 'indistinguishable' means the paired difference is inside two standard errors. It is not a tie broken in the model's favour.
- Closing lines come from football-data.co.uk domestic leagues only. Cup and European fixtures are absent by construction.