Calibration
The record says we do not beat the consensus, and we have not withdrawn it. This page tests the one claim that survives it. When the board says a player has a 6% chance to finish #1 at his position, does that happen about 6% of the time? Being honest about your own uncertainty is not the same thing as having an edge, which is exactly why we are allowed to ask it.
p 0.598
The dashed diagonal is a perfectly honest forecast: above it we promised too much, below it too little. Markers are sized by how many forecasts landed in the band — near-uniform here, because the bands are cut to hold equal numbers.
Every active player on every weekly board, 5866 forecasts in all, against whether he really finished #1 at his position that week. Bands are cut by population rather than by width, because almost every player on a board has a near-zero chance and ten equal-width bins would put the whole season in the first one. The equal-width cut gives ECE 0.0015, so the choice is not carrying the answer.
| Band | Forecasts | Promised | Happened | Gap |
|---|---|---|---|---|
| 0.00%–0.01% | 1542 | 0.00% | 0.00% | +0.00pp |
| 0.01%–0.04% | 545 | 0.03% | 0.00% | +0.03pp |
| 0.04%–0.07% | 360 | 0.05% | 0.28% | −0.23pp |
| 0.07%–0.14% | 606 | 0.10% | 0.00% | +0.10pp |
| 0.14%–0.29% | 506 | 0.20% | 0.20% | +0.00pp |
| 0.29%–0.87% | 563 | 0.53% | 0.71% | −0.18pp |
| 0.87%–1.94% | 580 | 1.36% | 1.90% | −0.54pp |
| 1.94%–3.84% | 579 | 2.81% | 2.42% | +0.39pp |
| 3.84% and up | 585 | 6.67% | 5.64% | +1.03pp |
The aggregate verdict is calibrated (p 0.598), but the aggregate is flattered by the bottom band: 1542 of the 5866 forecasts sit under 0.01%, where being right is free. The band that carries the board is the last row — 585 forecasts promising 6.67% and delivering 5.64%, an overstatement of +1.03pp at the end a reader would actually notice.
A forecast that says “the base rate” every single time is perfectly calibrated and completely useless, so calibration on its own is not a result. Splitting the Brier score separates the two: honesty is one term, discrimination is another.
| Component | Value | Reading |
|---|---|---|
| Brier | 0.01033 | mean squared error of the probabilities |
| Reliability | 0.00002 | the honesty term — lower is better |
| Resolution | 0.00030 | the discrimination term — higher is better |
| Uncertainty | 0.01079 | the base rate’s own variance, which nobody can beat |
| Skill | 0.0429 | share of that variance actually removed |
Skill of 0.0429 means the board removes 4.3% of the variance that a forecaster who only knew the base rate would face. It is a small number, and it is the honest one: we are well calibrated and only slightly informative, which is a different sentence from the one most projection sites print.
The dashed diagonal is a perfectly honest forecast: above it we promised too much, below it too little. Markers are sized by how many forecasts landed in the band — near-uniform here, because the bands are cut to hold equal numbers.
Sleeper publishes points, not probabilities. To ask both sides the same question their point projection is pushed through the identical uncertainty model as ours — same simulator, same per-position variance, same random stream, same pool of players — and the event is the winner inside that pool. The consensus probability here is ours, built from their numbers. It is not something they published, and nothing on this page should be read as a claim about what they say.
| Line | ECE | Brier | Reliability | Resolution | Skill |
|---|---|---|---|---|---|
| Ours | 0.0061 | 0.02224 | 0.00005 | 0.00072 | 0.0384 |
| Consensus | 0.0076 | 0.02229 | 0.00010 | 0.00077 | 0.0360 |
Position-weeks won 31–37, McNemar p 0.545; paired t on the per-week Brier difference p 0.533. The two lines are indistinguishable. That is the finding, and it is neither a win nor a retraction: the outright #1 is the tail of the ordering problem, so a model can be clearly worse at ordering — ours is — and still look level here.
The probabilities at one position in one week sum to 1, so the board is not really claiming a rate. It is claiming the winner is somebody on this list. Across 68 position-weeks it promised 68.1 winners and 64 arrived.
The actual #1 was on the board in 63 of 68. The 5 weeks he was not are almost the whole of the gap — so the miscalibration is not a shading of the probabilities. It is a hole in the candidate pool, which is a fixable thing rather than a philosophical one.
| Position | Forecasts | ECE | Bias | Our Brier | Crowd Brier |
|---|---|---|---|---|---|
| QB | 724 | 0.0080 | +0.14pp | 0.03501 | 0.03546 |
| RB | 1524 | 0.0048 | +0.07pp | 0.01914 | 0.01964 |
| WR | 2307 | 0.0031 | +0.09pp | 0.01423 | 0.01394 |
| TE | 1311 | 0.0061 | +0.00pp | 0.03458 | 0.03432 |
Pooled and per-position figures are cut on different bins, each quantile-binned on its own distribution, so they do not have to agree in direction — and here they do not. Neither gap is significant and neither should be quoted alone.
- It is not an edge. Calibration is a claim about honesty. The record still says we do not beat the consensus and this page does not touch it.
- It is one season. Nothing here is a persistence claim. It has to run on more seasons before it carries any weight.
- The forecasts are not independent. Everyone at one position in one week is competing for a single winner, which makes the per-forecast p-value conservative. The position-week tests are the ones to lean on.
- The consensus line is constructed. It is Sleeper’s points through our uncertainty model, eligibility set at 5 of their PPR points so we cannot pick the players we like.