Every prediction the model has made, graded — wins and losses both, nothing filtered.
What these numbers describe
97% of these 32,286 predictions (31,364) are the walk-forward Elo floor, not the live model. The live model accounts for 922. Read the figures below as describing the floor.
The floor is Elo + Poisson with ratings frozen at each match — no news, no calibration, no market. It is a lower bound: the live model should beat it.
A full analysis adds news, edges and market context on top of the model; the floor described above is the model alone. Both were made before the outcome was known, which is why they can be scored together. Full analyses accumulate as fixtures are auto-analysed.
Aggregate sport accuracy hides huge per-league variance. A row only publishes once it has ≥ 30 graded samples — below that the estimate is too noisy to be honest. Sorted best first; mid-table is where the model is mostly random.
vs naiveis the Brier score against what an uninformed guess scores in that league's market — 0.667 where a draw is possible, 0.500 where it is not. Positive means the model beats guessing. This is a much lower bar than the closing market, which the model loses to by 10.8σ on football (see the market baseline). Both bars are shown because publishing only the one it fails is no more honest than publishing only the one it clears.
What we want: Brier trending down (we're more calibrated) and accuracy drifting up. A flat Brier with rising accuracy usually means we're getting sharper on the obvious calls but not fixing systematic overconfidence.
Lower is better. A descending curve = the model is becoming more calibrated as more data accumulates. The dashed line marks the 5% threshold for "well-calibrated". Sudden spikes can signal regime changes (rule changes, calibrator drift) worth investigating.
A well-calibrated model lands every dot on the diagonal. Above the diagonal = under-confident (we said 70%, it actually won 80% of the time); below = over-confident, the dangerous direction because confident wrong bets cost the most.
Same data as above, broken out per bucket. Useful for spotting which confidence range is off — the reliability diagram shows direction, this shows magnitude per bin.