Skip to content
No picks tracked yetSave a pick and this strip becomes your running P/L.See tomorrow's slate →
betsetgo
Track record

Model vs the closing line

Every other number on this site scores the model against itself — against the outcomes it predicted, or against a fair price derived from its own probability. This page scores it against someone else's opinion: the closing market on the same matches. It is the only measurement here that predicts whether betting the model makes money, and the only one whose answer can be no.

How this was measured: model probabilities recorded before kickoff, scored against the devigged closing line on the same matches, paired per match so the two sides never describe different samples. Brier and log loss with standard errors; no cherry-picking, no post-hoc filtering, and the result is published whichever way it falls. Built and measured by nvision-data.

Football

Big-five leagues · closing prices from football-data.co.uk

Walk-forward floor (Elo only)

3,507 matches · 2024-08-15 → 2026-07-23

Market beats the model

Over 3,507 matches the model's Brier is 0.0272 WORSE than the closing line's (± 0.0025). This is the expected result: the close is a strong forecast. It means flat betting this model at market prices loses.

Brier scorelower is better
Model
0.6022
Closing line
0.5750
Difference
+0.0272market better
Log losslower is better
Model
1.0085
Closing line
0.9670
Difference
+0.0415market better
Argmax accuracyhigher is better
Model
50.9%
Closing line
54.3%
Difference
−3.5ppmarket better

Is the difference real?

+0.0272 ± 0.0025 (10.8σ)

Beyond 2 standard errors, so the gap is not sampling noise at this sample size.

Model closer to the truth

36.2%

Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.

Elo-only with ratings frozen as of each match — no calibration, no news, no market. The lowest defensible version of the model, so a loss here sizes the gap the live signal stack would have to close rather than condemning it.

Live model (Platt + news)

187 matches · 2026-04-26 → 2026-07-23

Not enough matches to say

187 matches with both a model probability and a closing line — below the 200 needed before a comparison means anything. No verdict.

Brier scorelower is better
Model
0.6218
Closing line
0.6024
Difference
+0.0194market better
Log losslower is better
Model
1.0430
Closing line
1.0061
Difference
+0.0369market better
Argmax accuracyhigher is better
Model
50.3%
Closing line
50.8%
Difference
−0.5ppmarket better

Is the difference real?

+0.0194 ± 0.0157 (1.2σ)

Below the 200-match floor for publishing a verdict, so no claim either way.

Model closer to the truth

48.7%

Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.

The probabilities the site actually published: Elo plus Platt calibration and news features. Market-blind — the market snapshot only ever enters the edge calculation, never the probability.

Split by how much the model knew

Elo only becomes informative once it has seen a team play. Where it has, the figures above are a measurement of the model. Where it hasn't, the model emits nearly the same probabilities for every fixture, and scoring those against the close measures our prior rather than our model. Both are shown; neither replaces the headline.

Live model

Elo had enough games

+0.0139 ± 0.0188 (0.7σ)

117 matches · model closer on 52.1%

PD 39 · SA 24 · PL 20 · FL1 20 · BL1 14

Elo was still near its prior

+0.0284 ± 0.0278 (1.0σ)

70 matches · model closer on 42.9%

BSA 49 · PPL 19 · DED 2

Walk-forward floor

Elo had enough games

+0.0271 ± 0.0025 (10.8σ)

3,440 matches · model closer on 36.3%

PD 750 · PL 744 · SA 735 · FL1 608 · BL1 602 · ELC 1

Elo was still near its prior

+0.0320 ± 0.0276 (1.2σ)

67 matches · model closer on 34.3%

BSA 47 · PPL 16 · DED 4

Where these prices came from

Closing lines held
3,550
Matches compared
3,513
Excluded
37
Mean overround
3.88%

Price source

  • fd-co-uk:pinnacle-close2,631 matches
  • fd-co-uk:market-avg-close882 matches

Why 37 were excluded

No model prediction
37
Analysis after kickoff
1

Tennis

ATP + WTA · closing prices from tennis-data.co.uk

Read the difference, not the scores

Tennis results are stored with the winner in the home slot, so both columns below answer “what probability did you give the player who went on to win?” Favourites win more often than not, so both Brier scores come out lower — better-looking — than football's, and that is an artefact of the labelling, not tennis skill. Only the model-minus-market difference is comparable between the two sports.

The comparison is still fair: model and market are scored on the same quantity over the same matches, and neither knew the result. Tennis Elo carries no home-court bonus, so nothing leaks into the slot that always holds the winner.

Walk-forward floor (Elo only)

8,267 matches · 2024-12-29 → 2026-07-19

Market beats the model

Over 8,267 matches the model's Brier is 0.0307 WORSE than the closing line's (± 0.0023). This is the expected result: the close is a strong forecast. It means flat betting this model at market prices loses.

Brier scorelower is better
Model
0.4387
Closing line
0.4079
Difference
+0.0307market better
Log losslower is better
Model
0.6277
Closing line
0.5925
Difference
+0.0353market better
Argmax accuracyhigher is better
Model
64.2%
Closing line
69.3%
Difference
−5.1ppmarket better

Is the difference real?

+0.0307 ± 0.0023 (13.5σ)

Beyond 2 standard errors, so the gap is not sampling noise at this sample size.

Model closer to the truth

43.8%

Share of matches where the model put more probability on the outcome that actually happened than the market did. 50% is a coin flip against the close.

Elo-only with ratings frozen as of each match — no calibration, no news, no market. The lowest defensible version of the model, so a loss here sizes the gap the live signal stack would have to close rather than condemning it.

Where these prices came from

Matches paired
8,273
Matches compared
8,267
Unusable price
6
Mean overround
5.22%

Bet365 closing prices published by tennis-data.co.uk, joined to our own player records. Until these landed the sport had no second opinion at all — the previous data source carried results and no odds, so tennis was the last market here that had never been tested.

Only the walk-forward floor is measured here, where football above also shows the live model. That is not a result — tennis predictions come from the backfill simulation, and the live stack has not yet published graded probabilities on matches these prices cover. When it has, a second card will appear rather than this one changing.

Reading this honestly

  • Lower is better for both Brier and log loss — they measure how wrong a forecast was. The side with the smaller number is the better forecaster.
  • The market is expected to win. A devigged closing price from a sharp book aggregates far more information than an Elo model. Beating it is rare, and a claim to have done so deserves more scepticism than a claim not to have.
  • The verdict is decided on Brier, not log loss. Both are proper scores, but log loss is unbounded — one confident miss can move its mean more than the rest of the sample combined. Brier is bounded, so its paired standard error is a trustworthy noise floor. Both are shown so you can disagree.

Football specifics

  • The market is devigged multiplicatively, which slightly understates the favourite. Sharpest-book columns are preferred to keep that bias small.
  • A season can mix price sources. football-data.co.uk stopped filling Pinnacle closing odds in January 2026, so later matches fall back to the market average — measured overround 6.2% against Pinnacle's 3.0%. A wider-margin book is a slightly weaker benchmark, and the bookmaker list shows the split per season.
  • Both prediction sources are market-blind: walk-forward is Elo-only with frozen ratings, and the live model's probabilities never read the market snapshot.
  • A verdict of 'indistinguishable' means the paired difference is inside two standard errors. It is not a tie broken in the model's favour.
  • Closing lines come from football-data.co.uk domestic leagues only. Cup and European fixtures are absent by construction.

Tennis specifics

  • Tennis rows store the winner in the home slot, so both sides answer 'what probability did you give the eventual winner?'. Favourites usually win, so both Brier scores look better than football's for reasons of labelling, not skill. Only the model-minus-market difference is comparable across sports.
  • The comparison itself is fair: model and market are scored on the same quantity over the same matches, and neither was computed with knowledge of the result. Tennis Elo carries no home advantage, so no bonus leaks into the slot that always holds the winner.
  • Predictions are the walk-forward Elo floor with ratings frozen as of each match — no calibration, no news, no market input.
  • Prices are Bet365 closing odds published by tennis-data.co.uk, joined to our own player records.