Calibration

Season 2025 · 5866 forecasts
weeks 218 · 4000 sims a position-week

The record says we do not beat the consensus, and we have not withdrawn it. This page tests the one claim that survives it. When the board says a player has a 6% chance to finish #1 at his position, does that happen about 6% of the time? Being honest about your own uncertainty is not the same thing as having an edge, which is exactly why we are allowed to ask it.

The board we publish5866 forecasts · base rate 1.09%
calibratedECE 0.0024
p 0.598
Reliability curve for the 2025 board: forecast probability against observed frequency0%0%2.5%2.5%5.0%5.0%7.5%7.5%10.0%10.0%0.00%–0.01% band, 1542 forecasts: promised 0.00%, happened 0.00%0.01%–0.04% band, 545 forecasts: promised 0.03%, happened 0.00%0.04%–0.07% band, 360 forecasts: promised 0.05%, happened 0.28%0.07%–0.14% band, 606 forecasts: promised 0.10%, happened 0.00%0.14%–0.29% band, 506 forecasts: promised 0.20%, happened 0.20%0.29%–0.87% band, 563 forecasts: promised 0.53%, happened 0.71%0.87%–1.94% band, 580 forecasts: promised 1.36%, happened 1.90%1.94%–3.84% band, 579 forecasts: promised 2.81%, happened 2.42%3.84% and up band, 585 forecasts: promised 6.67%, happened 5.64%WHAT WE PROMISEDWHAT HAPPENED
Top Performer AI board
The dashed diagonal is a perfectly honest forecast: above it we promised too much, below it too little. Markers are sized by how many forecasts landed in the band — near-uniform here, because the bands are cut to hold equal numbers.

Every active player on every weekly board, 5866 forecasts in all, against whether he really finished #1 at his position that week. Bands are cut by population rather than by width, because almost every player on a board has a near-zero chance and ten equal-width bins would put the whole season in the first one. The equal-width cut gives ECE 0.0015, so the choice is not carrying the answer.

BandForecastsPromisedHappenedGap
0.00%–0.01%15420.00%0.00%+0.00pp
0.01%–0.04%5450.03%0.00%+0.03pp
0.04%–0.07%3600.05%0.28%−0.23pp
0.07%–0.14%6060.10%0.00%+0.10pp
0.14%–0.29%5060.20%0.20%+0.00pp
0.29%–0.87%5630.53%0.71%−0.18pp
0.87%–1.94%5801.36%1.90%−0.54pp
1.94%–3.84%5792.81%2.42%+0.39pp
3.84% and up5856.67%5.64%+1.03pp

The aggregate verdict is calibrated (p 0.598), but the aggregate is flattered by the bottom band: 1542 of the 5866 forecasts sit under 0.01%, where being right is free. The band that carries the board is the last row — 585 forecasts promising 6.67% and delivering 5.64%, an overstatement of +1.03pp at the end a reader would actually notice.

Honest, or just vague?

A forecast that says “the base rate” every single time is perfectly calibrated and completely useless, so calibration on its own is not a result. Splitting the Brier score separates the two: honesty is one term, discrimination is another.

ComponentValueReading
Brier0.01033mean squared error of the probabilities
Reliability0.00002the honesty term — lower is better
Resolution0.00030the discrimination term — higher is better
Uncertainty0.01079the base rate’s own variance, which nobody can beat
Skill0.0429share of that variance actually removed

Skill of 0.0429 means the board removes 4.3% of the variance that a forecaster who only knew the base rate would face. It is a small number, and it is the honest one: we are well calibrated and only slightly informative, which is a different sentence from the one most projection sites print.

Ours against the consensus2913 paired · 68 position-weeks
Reliability curves for our projections and the consensus, on the same pool of players0%0%2.5%2.5%5.0%5.0%7.5%7.5%10.0%10.0%0.00%–0.01% band, 299 forecasts: promised 0.00%, happened 0.00%0.01%–0.11% band, 291 forecasts: promised 0.05%, happened 0.34%0.11%–0.34% band, 298 forecasts: promised 0.21%, happened 0.67%0.34%–0.74% band, 280 forecasts: promised 0.53%, happened 1.43%0.74%–1.26% band, 295 forecasts: promised 1.01%, happened 0.68%1.26%–1.99% band, 292 forecasts: promised 1.62%, happened 2.74%1.99%–2.89% band, 288 forecasts: promised 2.42%, happened 1.74%2.89%–4.09% band, 287 forecasts: promised 3.44%, happened 2.79%4.09%–6.29% band, 292 forecasts: promised 5.00%, happened 3.77%6.29% and up band, 291 forecasts: promised 9.13%, happened 9.62%0.00%–0.01% band, 165 forecasts: promised 0.00%, happened 0.61%0.01%–0.11% band, 288 forecasts: promised 0.05%, happened 0.00%0.11%–0.34% band, 305 forecasts: promised 0.21%, happened 0.33%0.34%–0.74% band, 324 forecasts: promised 0.52%, happened 0.93%0.74%–1.26% band, 267 forecasts: promised 0.97%, happened 1.12%1.26%–1.99% band, 320 forecasts: promised 1.62%, happened 0.31%1.99%–2.89% band, 362 forecasts: promised 2.41%, happened 2.76%2.89%–4.09% band, 345 forecasts: promised 3.44%, happened 4.93%4.09%–6.29% band, 283 forecasts: promised 5.02%, happened 2.83%6.29% and up band, 254 forecasts: promised 9.03%, happened 9.84%WHAT WE PROMISEDWHAT HAPPENED
Top Performer AIConsensus (Sleeper points)
The dashed diagonal is a perfectly honest forecast: above it we promised too much, below it too little. Markers are sized by how many forecasts landed in the band — near-uniform here, because the bands are cut to hold equal numbers.

Sleeper publishes points, not probabilities. To ask both sides the same question their point projection is pushed through the identical uncertainty model as ours — same simulator, same per-position variance, same random stream, same pool of players — and the event is the winner inside that pool. The consensus probability here is ours, built from their numbers. It is not something they published, and nothing on this page should be read as a claim about what they say.

LineECEBrierReliabilityResolutionSkill
Ours0.00610.022240.000050.000720.0384
Consensus0.00760.022290.000100.000770.0360

Position-weeks won 3137, McNemar p 0.545; paired t on the per-week Brier difference p 0.533. The two lines are indistinguishable. That is the finding, and it is neither a win nor a retraction: the outright #1 is the tail of the ordering problem, so a model can be clearly worse at ordering — ours is — and still look level here.

Where the shortfall really ispool coverage 92.6%

The probabilities at one position in one week sum to 1, so the board is not really claiming a rate. It is claiming the winner is somebody on this list. Across 68 position-weeks it promised 68.1 winners and 64 arrived.

The actual #1 was on the board in 63 of 68. The 5 weeks he was not are almost the whole of the gap — so the miscalibration is not a shading of the probabilities. It is a hole in the candidate pool, which is a fixable thing rather than a philosophical one.

PositionForecastsECEBiasOur BrierCrowd Brier
QB7240.0080+0.14pp0.035010.03546
RB15240.0048+0.07pp0.019140.01964
WR23070.0031+0.09pp0.014230.01394
TE13110.0061+0.00pp0.034580.03432

Pooled and per-position figures are cut on different bins, each quantile-binned on its own distribution, so they do not have to agree in direction — and here they do not. Neither gap is significant and neither should be quoted alone.

What this does not say
  • It is not an edge. Calibration is a claim about honesty. The record still says we do not beat the consensus and this page does not touch it.
  • It is one season. Nothing here is a persistence claim. It has to run on more seasons before it carries any weight.
  • The forecasts are not independent. Everyone at one position in one week is competing for a single winner, which makes the per-forecast p-value conservative. The position-week tests are the ones to lean on.
  • The consensus line is constructed. It is Sleeper’s points through our uncertainty model, eligibility set at 5 of their PPR points so we cannot pick the players we like.