Model vs the closing line
Every other number on this site scores the model against itself — against the outcomes it predicted, or against a fair price derived from its own probability. This page scores it against someone else's opinion: the closing market on the same matches. It is the only measurement here that predicts whether betting the model makes money, and the only one whose answer can be no.
How this was measured: model probabilities recorded before kickoff, scored against the devigged closing line on the same matches, paired per match so the two sides never describe different samples. Brier and log loss with standard errors; no cherry-picking, no post-hoc filtering, and the result is published whichever way it falls. Built and measured by nvision-data.
Football
Big-five leagues · closing prices from football-data.co.ukWalk-forward floor (Elo only)
3,507 matches · 2024-08-15 → 2026-07-23Market beats the model
Over 3,507 matches the model's Brier is 0.0272 WORSE than the closing line's (± 0.0025). This is the expected result: the close is a strong forecast. It means flat betting this model at market prices loses.
- Model
- 0.6022
- Closing line
- 0.5750
- Difference
- +0.0272market better
- Model
- 1.0085
- Closing line
- 0.9670
- Difference
- +0.0415market better
- Model
- 50.9%
- Closing line
- 54.3%
- Difference
- −3.5ppmarket better
Is the difference real?
+0.0272 ± 0.0025 (10.8σ)
Beyond 2 standard errors, so the gap is not sampling noise at this sample size.
Model closer to the truth
36.2%
Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.
Elo-only with ratings frozen as of each match — no calibration, no news, no market. The lowest defensible version of the model, so a loss here sizes the gap the live signal stack would have to close rather than condemning it.
Live model (Platt + news)
187 matches · 2026-04-26 → 2026-07-23Not enough matches to say
187 matches with both a model probability and a closing line — below the 200 needed before a comparison means anything. No verdict.
- Model
- 0.6218
- Closing line
- 0.6024
- Difference
- +0.0194market better
- Model
- 1.0430
- Closing line
- 1.0061
- Difference
- +0.0369market better
- Model
- 50.3%
- Closing line
- 50.8%
- Difference
- −0.5ppmarket better
Is the difference real?
+0.0194 ± 0.0157 (1.2σ)
Below the 200-match floor for publishing a verdict, so no claim either way.
Model closer to the truth
48.7%
Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.
The probabilities the site actually published: Elo plus Platt calibration and news features. Market-blind — the market snapshot only ever enters the edge calculation, never the probability.
Split by how much the model knew
Elo only becomes informative once it has seen a team play. Where it has, the figures above are a measurement of the model. Where it hasn't, the model emits nearly the same probabilities for every fixture, and scoring those against the close measures our prior rather than our model. Both are shown; neither replaces the headline.
Live model
Elo had enough games
+0.0139 ± 0.0188 (0.7σ)
117 matches · model closer on 52.1%
PD 39 · SA 24 · PL 20 · FL1 20 · BL1 14
Elo was still near its prior
+0.0284 ± 0.0278 (1.0σ)
70 matches · model closer on 42.9%
BSA 49 · PPL 19 · DED 2
Walk-forward floor
Elo had enough games
+0.0271 ± 0.0025 (10.8σ)
3,440 matches · model closer on 36.3%
PD 750 · PL 744 · SA 735 · FL1 608 · BL1 602 · ELC 1
Elo was still near its prior
+0.0320 ± 0.0276 (1.2σ)
67 matches · model closer on 34.3%
BSA 47 · PPL 16 · DED 4
Where these prices came from
- Closing lines held
- 3,550
- Matches compared
- 3,513
- Excluded
- 37
- Mean overround
- 3.88%
Price source
- fd-co-uk:pinnacle-close2,631 matches
- fd-co-uk:market-avg-close882 matches
Why 37 were excluded
- No model prediction
- 37
- Analysis after kickoff
- 1
Tennis
ATP + WTA · closing prices from tennis-data.co.ukRead the difference, not the scores
Tennis results are stored with the winner in the home slot, so both columns below answer “what probability did you give the player who went on to win?” Favourites win more often than not, so both Brier scores come out lower — better-looking — than football's, and that is an artefact of the labelling, not tennis skill. Only the model-minus-market difference is comparable between the two sports.
The comparison is still fair: model and market are scored on the same quantity over the same matches, and neither knew the result. Tennis Elo carries no home-court bonus, so nothing leaks into the slot that always holds the winner.
Walk-forward floor (Elo only)
8,267 matches · 2024-12-29 → 2026-07-19Market beats the model
Over 8,267 matches the model's Brier is 0.0307 WORSE than the closing line's (± 0.0023). This is the expected result: the close is a strong forecast. It means flat betting this model at market prices loses.
- Model
- 0.4387
- Closing line
- 0.4079
- Difference
- +0.0307market better
- Model
- 0.6277
- Closing line
- 0.5925
- Difference
- +0.0353market better
- Model
- 64.2%
- Closing line
- 69.3%
- Difference
- −5.1ppmarket better
Is the difference real?
+0.0307 ± 0.0023 (13.5σ)
Beyond 2 standard errors, so the gap is not sampling noise at this sample size.
Model closer to the truth
43.8%
Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.
Elo-only with ratings frozen as of each match — no calibration, no news, no market. The lowest defensible version of the model, so a loss here sizes the gap the live signal stack would have to close rather than condemning it.
Where these prices came from
- Matches paired
- 8,273
- Matches compared
- 8,267
- Unusable price
- 6
- Mean overround
- 5.22%
Bet365 closing prices published by tennis-data.co.uk, joined to our own player records. Until these landed the sport had no second opinion at all — the previous data source carried results and no odds, so tennis was the last market here that had never been tested.
Only the walk-forward floor is measured here, where football above also shows the live model. That is not a result — tennis predictions come from the backfill simulation, and the live stack has not yet published graded probabilities on matches these prices cover. When it has, a second card will appear rather than this one changing.
Reading this honestly
- Lower is better for both Brier and log loss — they measure how wrong a forecast was. The side with the smaller number is the better forecaster.
- The market is expected to win. A devigged closing price from a sharp book aggregates far more information than an Elo model. Beating it is rare, and a claim to have done so deserves more scepticism than a claim not to have.
- The verdict is decided on Brier, not log loss. Both are proper scores, but log loss is unbounded — one confident miss can move its mean more than the rest of the sample combined. Brier is bounded, so its paired standard error is a trustworthy noise floor. Both are shown so you can disagree.
Football specifics
- The market is devigged multiplicatively, which slightly understates the favourite. Sharpest-book columns are preferred to keep that bias small.
- A season can mix price sources. football-data.co.uk stopped filling Pinnacle closing odds in January 2026, so later matches fall back to the market average — measured overround 6.2% against Pinnacle's 3.0%. A wider-margin book is a slightly weaker benchmark, and the bookmaker list shows the split per season.
- Both prediction sources are market-blind: walk-forward is Elo-only with frozen ratings, and the live model's probabilities never read the market snapshot.
- A verdict of 'indistinguishable' means the paired difference is inside two standard errors. It is not a tie broken in the model's favour.
- Closing lines come from football-data.co.uk domestic leagues only. Cup and European fixtures are absent by construction.
Tennis specifics
- Tennis rows store the winner in the home slot, so both sides answer 'what probability did you give the eventual winner?'. Favourites usually win, so both Brier scores look better than football's for reasons of labelling, not skill. Only the model-minus-market difference is comparable across sports.
- The comparison itself is fair: model and market are scored on the same quantity over the same matches, and neither was computed with knowledge of the result. Tennis Elo carries no home advantage, so no bonus leaks into the slot that always holds the winner.
- Predictions are the walk-forward Elo floor with ratings frozen as of each match — no calibration, no news, no market input.
- Prices are Bet365 closing odds published by tennis-data.co.uk, joined to our own player records.