What doesn’t.work

17 claims tested · 12 refuted
updated 2026-08-30

Fantasy football runs on numbers nobody has checked. These are the ones we checked. Each entry names the claim as it is actually made, the test, the sample, and what we changed in our own product because of the answer — a finding that changed nothing was never believed. Our own claims are in this list, judged by the same rule, and one of them is refuted.

Industry claimtested 2026-08-29
REFUTED
Start your player against the defenses that give up the most points to his position.

At quarterback and receiver, how many points a defense allowed to a position last season tells you essentially nothing about this season.

Evidence

  • Year-over-year persistence r = 0.08 at QB and r = 0.05 at WR, against a significance threshold of 0.20
  • RB (r = 0.31) and TE (r = 0.25) are real but weak
  • Measured on 96 team-season transitions, drift-adjusted so a league-wide trend cannot be mistaken for a team

Sample

32 teams over four seasons, 2022-2025, 96 transitions

Method

Correlate a defense's PPR allowed per game to a position in season Y against the same figure in Y+1, after removing league drift. Compare the predictive error of carrying the number forward against simply using the league average.

What we changed

Points-allowed facts can no longer carry an EDGE bullet in any brief. Our own Week 1 quarterback brief cited this stat three times before the gate landed; it is now demoted with the reason recorded.

Caveat

This measures the season-to-season persistence of the stat, not whether a defense is weak in the week you are looking at. A within-season signal may exist; we have not tested it.

Working: docs/research/2026-08-29-persistence-gate.md
Reproduce: npx tsx scripts/carryover.ts

Industry claimtested 2026-08-29
REFUTED
This team ran play-action 28% of the time last year, so expect that again.

Most scheme tendencies do not carry from one season to the next, and quoting last season's rate raw is worse than saying nothing about the team at all.

Evidence

  • Only pre-snap motion persists strongly (r = 0.74); shotgun 0.59, blitz 0.57, red-zone rush 0.55, no-huddle 0.51
  • For play-action, deep rate, screen rate and the coverage families, a team's own prior number predicts next season WORSE than the league average does
  • Raw carry-forward loses to the league mean on 7 of 11 rates; shrinking toward the mean by measured persistence beats it on all 11

Sample

32 teams over four seasons, 2022-2025, 96 transitions

Method

Compute each team's rate per season from play-by-play, participation and FTN charting; correlate Y against Y+1 after removing league drift; compare mean absolute error of carry-forward, league mean, and a shrunk estimator.

What we changed

Every team rate in a brief is now the shrunk expectation stated beside the observation and the league average, and a rate whose persistence is indistinguishable from zero cannot carry an EDGE bullet at any sample size.

Working: docs/research/2026-08-29-scheme-carryover.md
Reproduce: npx tsx scripts/carryover.ts

Industry claimtested 2026-08-29
REFUTED
This receiver is elite against man coverage.

For most split types the claim carries no information beyond the player being good. Knowing his overall rate predicts next season's split better than the split itself does.

Evidence

  • Blitz splits: the split predicts next season at r = 0.854, but the player's overall rate predicts it at r = 0.876
  • Game-script splits: 0.837 against 0.866, the same way round
  • Both come out negative against the player's own baseline; only target depth carries information beyond the player, at +6.5%

Sample

785 to 1,173 player-season pairs depending on split type, players with at least 50 opportunities in both seasons

Method

For each player and split bucket with n >= 15 in consecutive seasons, test whether the split value predicts next season's split better than the player's overall PPR per opportunity does.

What we changed

Split facts are gated on beating the player's own baseline, and the player's overall rate was promoted to a first-class fact — it had never been published at all. Citable facts per brief fell from 47, 31, 43 and 37 to 8, 3, 10 and 6.

Working: docs/research/2026-08-29-player-split-persistence.md
Reproduce: npx tsx scripts/player-persistence.ts

Industry claimtested 2026-08-29
UNMEASURABLE
Public charting data can be compared season to season.

The nflverse participation coverage labels for 2023 and 2024 disagree with external charting badly enough that no cross-season coverage comparison can be trusted.

Evidence

  • League man-coverage rate reads 28.8%, 41.3%, 48.4%, 30.9% across 2022-2025 — defenses do not move like that
  • Cover-1 and Cover-3 swing 15 percentage points in exact mirror image, the signature of a classifier moving the boundary between two lookalike single-high shells
  • Published 2023 figures put Cover-1 near 18.7% and Cover-3 near 35.7%; our feed reads 34.8% and 19.9% — the two families near-inverted, against a documented league trend toward zone

Sample

Four seasons of scrimmage dropbacks, roughly 19,000 per season, participation join rate 100% throughout

Method

Compare league-level coverage distributions across seasons on a fixed denominator, check feed schema changes and label completeness, and compare against published external charting.

What we changed

2023 and 2024 are quarantined in code. Every participation-derived fact is demoted when a quarantined season is in the pool, which also means we cannot state whether defensive tendencies persist at all — there is no consecutive trusted season pair left.

Caveat

Charting providers define coverage families differently, so the external comparison is corroboration rather than proof. It does not need to be proof to stop us citing the numbers.

Working: docs/research/2026-08-29-coverage-label-quarantine.md
Reproduce: npx tsx scripts/carryover.ts

Our own claimtested 2026-08-29
REFUTED
Our model beats the free consensus at picking the week's top performer.

We do not beat it. We name the actual week-winner about as often, and we finish further back when we miss.

Evidence

  • 136 blind paired picks: our call finished top-3 32 times against the consensus's 40
  • Outright #1s tied 12-12; average finish 16.14 against 12.54
  • All four configurations of the engine lose on the blind seasons, and the one chosen beforehand on principle turned out to be the worst of the four

Sample

2022 and 2023, 17 weeks x 4 positions x 2 seasons = 136 paired picks, neither season run before the rule was fixed

Method

Pre-register the configuration, endpoint, test and labels in a committed file; run the blind seasons once; report McNemar's exact test on podium rate and paired tests on finish.

What we changed

Every call on the board carries the verdict stamped on its face, and the earlier claim of an edge at running back was withdrawn in public rather than quietly deleted.

Caveat

136 paired picks can only detect a large effect. This is not proof that no edge exists; it is proof that none has been demonstrated.

Working: docs/receipts/validation-four-season.md
Reproduce: npx tsx scripts/backtest.ts 2022 --opp-on && npx tsx scripts/pooled.ts

Our own claimtested 2026-08-30
SUPPORTED
Our win probabilities mean what they say.

When the board says a player has a 6% chance to finish #1 at his position, it happens about 6% of the time.

Evidence

  • Expected calibration error 0.0024 over 5,866 forecasts, base rate 1.09%
  • Bias +0.0007 — we promise very slightly more winners than arrive
  • Paired t on the residuals gives t = 0.53, p = 0.598, so the gap from honesty is not distinguishable from zero
  • Equal-width bins instead of population bins give ECE 0.0015, so the binning choice is not carrying the result

Sample

5,866 forecasts, one per active player per position-week, 2025 weeks 2-18

Method

Bin the published pTop by population, compare predicted frequency against observed frequency, and decompose the Brier score into reliability, resolution and uncertainty.

What we changed

This is now the one accuracy-adjacent claim the site is allowed to make, and it is published at /calibration with the reliability curve. It is deliberately NOT a claim to beat the consensus, which the same repository refutes.

Caveat

Calibration is not accuracy. A forecast that says 1% for everybody would also be near-calibrated on a 1% base rate; the resolution component (0.00030) is what says we are doing more than that, and it is small.

Working: docs/research/2026-08-30-calibration-2025.md
Reproduce: npm run calibration -- 2025

Industry claimtested 2026-08-30
UNMEASURABLE
The NFL gives away player tracking data, so anyone can build Next Gen Stats.

The data is real and free, but it is Next Gen Stats under a confidentiality clause — it can referee our own numbers privately and can never appear in anything we publish.

Evidence

  • Kaggle API returns 401 without a token; the CLI refuses before sending a request
  • One 2017 game is genuinely public on the NFL’s own GitHub: 316,025 rows, 23 tracked objects, 10 Hz, verified by downloading it
  • The 2026 release covers 2023 and 2024 — exactly the two seasons whose coverage labels we quarantined — with a coverage call from a different charting source

Sample

Eight competition releases surveyed; one 2017 game opened and measured directly

Method

Probe every documented endpoint with curl and record real HTTP status codes; download and count the one public file rather than citing its documentation.

What we changed

Tracking data is scoped as an internal auditor for the quarantined coverage labels, never as a published number, and no tracking file may enter this repository. A Kaggle token is a human step that has not been taken.

Caveat

The licence terms were read from pages that resist automated fetching, so the quoted clauses should be re-checked in a browser before any commercial decision rests on them.

Working: docs/research/2026-08-30-big-data-bowl-access.md
Reproduce: npm run bdb-probe

Industry claimtested 2026-08-30
REFUTED
Read last season's film to know how a defense will play this season.

A team's own current-season tape beats last season's more often than not, and for some tendencies neither beats simply knowing the league average.

Evidence

  • Across 32 testable metric-weeks in 2025: current season best 15 times, prior season 6, league mean 11
  • Blitz: current-season tape wins 10 of 16 weeks and every week from 12 onward bar one (week 18: 6.84pp error against the prior season's 8.43pp)
  • Pressure: no crossover exists — the league mean wins 7 weeks, and the spread between all three predictors is under a point most weeks

Sample

32 defences, 2025 weeks 3-18, two tendencies whose feed is comparable across the season boundary

Method

For each week, predict a defence's rate from (a) its current season to date, (b) all of its prior season, (c) the league to date, and compare mean absolute error against what it actually did that week.

What we changed

The weekly loop had hardcoded prior-season film for all eighteen weeks, with a comment admitting the crossover was unmeasured. That rule is retired: the film season is a per-tendency, week-aware decision.

Caveat

Only blitz and pressure could be tested. Coverage, man/zone and box — the tendencies a film room most wants — cannot cross this boundary because 2024's participation labels are quarantined. Enough to retire the blanket rule, not enough to claim a crossover week for anything else.

Working: docs/research/2026-08-30-film-season-crossover-2025.md
Reproduce: npx tsx scripts/crossover-study.ts 2025

Industry claimtested 2026-08-30
REFUTED
He is injury-prone / he is durable — discount him a couple of games.

How many games a player played last season tells you almost nothing about how many he plays next season, and at running back the league-wide group mean predicts better than his own number does.

Evidence

  • Correlation between games played in year Y and Y+1: r = 0.05 (RB), 0.10 (QB), 0.17 (WR), 0.29 (TE), against the rCrit = 0.20 bar this project uses elsewhere
  • Carry-forward mean absolute error LOSES to the group mean at running back (3.42 games against 3.36)
  • Nobody plays a full season either: players who finished top-12 at their position averaged 14.2 games the next year and only about half reached 16

Sample

624 player-seasons, 2021-2025, players finishing inside the drafted universe at their position

Method

Rank every player within position by PPR total in season Y; count regular-season games played in Y+1; correlate, and compare the predictive error of carrying a player's own games-played forward against simply using the group mean.

What we changed

Expected games in the season projection come from the position and the projected tier only. No player is ever docked games for a reputation. A currently reported condition (IR, PUP, a repaired ligament on today's roster) is a fact about now rather than a durability prediction, so it is applied separately, shown on the card, and vetoes a BUY.

Caveat

This measures season-to-season persistence of games played, which mixes injury with losing a job. It is the right quantity for a draft projection and the wrong one for a medical question.

Working: docs/receipts/draft-board-2026.md
Reproduce: npx tsx scripts/draft-backtest.ts 2023 2024 2025

Our own claimtested 2026-08-30
WEAK
Our season projection is better than just using last year's numbers.

Against the two baselines every drafter already has, the board wins 3-0 at running back and 2-1 at receiver, and establishes no edge at quarterback (2-1 but negative in 2025) or tight end (1-2). It is a real tool at two positions of four, which is why it is WEAK and not SUPPORTED, and the product says so before it shows a ranking.

Evidence

  • Pooled rank correlation RB 0.408 against 0.287 for the better baseline; the margin widens each season (0.26, 0.42, 0.55)
  • WR 0.338 against 0.300; beat both baselines in 2024 and 2025, tied in 2023
  • QB 0.097 and TE 0.050 — near zero for every method tested, ours included
  • Top-12 hit rate is 0.42-0.50 for OUR board and 0.42-0.50 for the baselines: nobody picks the eventual top twelve better than a coin flip

Sample

three preseason boards (2023, 2024, 2025) graded on the union of each method's top N and the actual top N

Method

Rebuild the board with no knowledge of the target season — no game logs, the season's first depth-chart snapshot rather than its newest, no betting lines — then Spearman against the realised finish, plus top-12 and top-24 hit rates.

What we changed

The product states the per-position verdict before it states a ranking. QB and TE boards are published as an opinion, not an edge. Mean absolute error is reported but carries no weight: it is biased by survivorship, because the players who finish top-24 are disproportionately the ones who stayed healthy.

Working: docs/receipts/draft-board-2026.md
Reproduce: npx tsx scripts/draft-backtest.ts 2023 2024 2025

Our own claimtested 2026-08-30
REFUTED
You cannot project a rookie, so leave him off the board until he plays.

Draft capital and opening depth-chart rank price a rookie's opportunity well enough to rank him, and the relationship is monotonic in both.

Evidence

  • First-round backs opening as the listed starter averaged 14.1 carries and 4.7 targets per team game in year one; sixth- and seventh-rounders behind two backs averaged 1.8 and 0.4
  • First-round receivers as the listed WR1 averaged 6.9 targets per team game against 0.06 for undrafted receivers behind two others
  • Measured per TEAM game, so games a rookie misses are already inside the rate rather than applied twice

Sample

2022-2025 rookie seasons at QB/RB/WR/TE, roster-joined to the season's opening depth chart

Method

Collapse each rookie's usage vector to one PPR-weighted scalar; fit a capital effect and a depth effect on the position's drafted-rookie base, each multiplier shrunk toward 1 by its own sample size.

What we changed

The weekly engine still sorts a no-game-log player below everyone it has observed, which is correct for a top-performer call. The draft board does not: it prices the role, because that is the entire question about a rookie.

Caveat

Cells are thin at the top of the board — two to six first-round backs across four seasons — which is exactly why the multipliers are shrunk rather than read raw.

Working: docs/receipts/draft-board-2026.md
Reproduce: npm run draft -- 2026

Industry claimtested 2026-08-30
REFUTED
Certain matchups widen or narrow a player's week-to-week production, even where they do not change his average.

Seven independent tests of the idea, and it does not survive any of them. The traits themselves persist; the widening they are supposed to cause does not exist as a team property at all.

Evidence

  • Blitz: +4.7% in quarterback residual SD per 1 SD of opponent blitz rate, p=0.130, and 0 of 13 tests survive BH FDR at q=0.10. Minnesota 2023 blitzed on 52.2% of dropbacks, the heaviest defence-season in four years, and NARROWED quarterback outcomes (residual SD 6.12 against a 7.12 league figure); the three lightest-blitzing defences all sat above league dispersion
  • The killer number: a defence's tendency to widen quarterback outcomes carries year to year at r=+0.026 across 96 transitions (p=0.80) against this project's rCrit=0.20 — while the blitz rate itself carries at r=0.576. A permutation test says there is no such defence trait to persist: between-defence variance 0.270 against a shuffled null of 0.248, p=0.234
  • Game script: Justin Jefferson's outcome SD was 9.78 as a 6+ point favourite and 9.62 as a 3+ point underdog, a ratio of 1.02. Across 12 spread and total specifications, 0 survive FDR and 0 reach even uncorrected p<0.05 in the 2025 holdout
  • Red-zone defence: a TE-only effect clears FDR in-sample and then fails out of sample, and the trait cannot persist anyway — the league's best red-zone defence fell to 13th, 18th, 11th and 16th the following year, four seasons running
  • Pace, coverage shell and funnel defences: all refuted, mean- and volume-controlled

Sample

2021-2025 play-by-play, participation and FTN charting; fit 2021-2024, held out on 2025

Method

Model log residual-variance with the player's own leave-one-out mean and weekly volume as covariates (and in the stricter form, player-season fixed effects), cluster-robust on defence-season, with Benjamini-Hochberg FDR across the whole test family and a permutation null.

What we changed

No matchup or boom/bust badge appears anywhere in the product, and no weekly brief may claim a defence makes a player more or less volatile.

Working: docs/receipts/matchup-variance-2026.md
Reproduce: see docs/receipts/matchup-variance-2026.md

Industry claimtested 2026-08-30
REFUTED
He is a boom-or-bust player — that is just who he is.

Week-to-week volatility barely persists once you remove the thing that causes it. Raw volatility looks like a strong, stable trait, and that is the player's scoring level wearing a costume.

Evidence

  • Controlled for mean, volume and games, year-over-year persistence of volatility is r=+0.086 at WR and +0.135 at RB — both under the rCrit=0.20 bar; TE +0.229 with a confidence interval straddling zero; only QB (+0.320) clears cleanly
  • UNCONTROLLED the same quantity looks compelling and would have shipped: RB r=+0.700, TE +0.572, QB +0.570, WR +0.470
  • Zack Moss was the most volatile qualifying back in the league in 2022 (100th percentile) on 5.0 points a game as a backup; given a real role in 2023 his volatility fell to the 76th. The boom/bust was his snap count, not his running style
  • Tank Bigsby, same shape: 100th percentile volatility in 2023 on 1.9 points a game, 87th in 2024 on a real role

Sample

625 player-seasons across four transitions, 2021-2025, with a 2024-25 holdout

Method

Coefficient of variation and log-SD residualised on the player-season mean, opportunity volume and games played, z-scored within position-season, drift-adjusted, correlated Y against Y+1.

What we changed

There is no boom/bust badge and no floor/ceiling label anywhere on the draft board. Where variance is discussed at all it is attributed to role and volume, which do persist, rather than to temperament.

Working: docs/receipts/matchup-variance-2026.md
Reproduce: see docs/receipts/matchup-variance-2026.md

Industry claimtested 2026-08-30
REFUTED
Look at red-zone and goal-line share — that is where the touchdowns, and the fantasy points, come from.

Red-zone share is a worse guide to next season than plain volume share, and where the two disagree it is the one that is wrong.

Evidence

  • Head to head on teammate pairs, 2021-2025: plain volume share ranked next season correctly 68.7% of the time (268/390), inside-the-ten share 62.1%, goal-to-go 63.7%, red-zone 'leverage' 54.4% — a coin flip
  • On points per game the gap is wider: RB volume 79.2% against 72.2% for inside-the-ten and 50.0% for leverage; WR volume 78.0% against 66.5% and 51.0%
  • It never beats what the board already knows from volume, and when it disagrees with volume it is backwards

Sample

390 teammate pairs at the same position, 2021-2025

Method

Compute inside-20, inside-10 and goal-to-go opportunity shares from play-by-play; rank teammate pairs by each metric and by plain volume share; score each against the pair's actual next-season finish.

What we changed

No red-zone or touchdown-equity statistic appears on a player card. The volume line does that job better.

Working: docs/receipts/player-metrics-2026.md
Reproduce: see docs/receipts/player-metrics-2026.md

Industry claimtested 2026-08-30
REFUTED
A defence that takes the ball away hands its offence more chances, so draft that offence's skill players.

Takeaways buy field position, not snaps, and none of it survives to the next August. The mechanism breaks at its third link and the driver does not carry.

Evidence

  • Within-team first differences: +1 takeaway per game buys +2.82 yards of starting field position (p=1.1e-12) and +0.593 offensive drives (p=3.8e-5) — but only +0.61 plays (p=0.40), because plays per drive falls 0.245 (p=0.0033). A short field means fewer snaps to finish the drive, and the two effects cancel
  • Persistence, drift-adjusted, n=128, against rCrit=0.20: takeaways/g r=+0.157, turnover margin +0.150, offensive drives +0.165, plays +0.158, and drive-start field position r=-0.039 — the one thing takeaways genuinely buy is the flattest series in the study
  • The lagged link a board would actually use is nil: last season's takeaways predict this season's drives at r=-0.008, plays +0.064, skill opportunities +0.06
  • Mean reversion is the strongest single result: r=-0.649 between a season's takeaway z and the change into the next (p=1.1e-16). Top-8 takeaway defences gave back 74% of their turnover margin
  • Jacksonville 2025, the case that prompted this: best starting field position in football (own 32.9, 1st) off the 2nd-most takeaways, and it produced the 13th-best fantasy offence. Drift-adjusted they gained +5.85 plays a game and the takeaway model explains 13% of them — 87% came from the offence, not the defence

Sample

160 team-seasons, 2021-2025, with a 2025 holdout; 61 tests, Benjamini-Hochberg at q=0.10

Method

Team-season panel of drives, plays, drive-start field position and takeaways from play-by-play; pooled, within-team first-difference and drift-adjusted persistence estimators, controlling for prior point differential and the preseason market line.

What we changed

No projection input, tier adjustment or edge bullet may key on takeaways, turnover margin, interception rate, fumble-recovery rate, or prior-season drive-start field position — in either direction, so no positive-regression note for low-takeaway teams either. The same-season field-position link may describe a WEEKLY matchup, where the rate is observed, and may never cross into the August board.

Caveat

The user's mechanism was half right and it is worth saying which half: the defence did buy Jacksonville the field position and the possessions. It did not buy the plays, and the scoring it did buy went to the quarterback rather than to anyone else on the roster.

Working: docs/receipts/regime-and-possessions-2026.md
Reproduce: see docs/receipts/regime-and-possessions-2026.md

Industry claimtested 2026-08-30
REFUTED
A new coaching staff shakes up who gets the ball, so a player's role is less predictable in year one.

A coaching change moves neither an offence's volume nor the spread of its players' role changes. What widens next season's error is a bad prior season, whoever is calling the plays.

Evidence

  • Zero of 18 mean tests on offensive volume and shape clear even raw p<0.05: plays/g +0.01 SD (p=0.98), drives/g +0.02 (p=0.51), pace +0.08 (p=0.78), pass rate -0.12
  • On 599 stayer player-seasons the standard-deviation ratio of change in opportunity share under a new staff is 1.040 (Levene p=0.21); target share 1.069, snap share 1.020, points per game 1.014 — no measurable scrambling
  • The only thing that reliably predicts a player's role change is mean reversion of his own prior share, slope -0.23, and it is identical under a new staff (-0.231) and a continuing one (-0.205)
  • The real driver of forecast error is a bad prior season: bad-prior teams that KEPT their staff carried their offensive profile at r=0.274 against 0.473 for good-prior continuity teams
  • One genuine year-one effect exists and is fantasy-irrelevant: a new head coach roughly doubles the dispersion of shotgun rate (variance ratio 2.15 raw, 1.94 quality-matched)

Sample

coaching changes 2022-2025 with a 2025 holdout; 599 stayer player-seasons

Method

Compare year-one change in team volume and in player opportunity share under new staffs against continuity teams, with Levene tests on dispersion and controls for prior point differential and the preseason market line.

What we changed

No card carries a coaching-change modifier in any direction — not a rank move, not a points adjustment, not a widened interval, and not a 'less certain' tag. Where a wider band is warranted it is keyed on a bad prior season instead. The staff change stays on the card as narrative only.

Caveat

Twenty-one of thirty-two offences changed hands for 2026. A tag that fires on two thirds of the league was never going to separate anything.

Working: docs/receipts/regime-and-possessions-2026.md
Reproduce: see docs/receipts/regime-and-possessions-2026.md

Our own claimtested 2026-08-30
UNMEASURABLE
The 2025 play-by-play defensive turnover feed is complete.

The New York Jets' 2025 defensive interceptions are missing from the feed. Every turnover-derived figure for that team is unusable.

Evidence

  • 0 interceptions logged for NYJ across 17 regular-season games and 1,419 defensive plays, in both play_by_play_2025.csv and stats_player_week_2025.csv
  • The league minimum excluding NYJ is 6, and 380 interceptions were thrown league-wide in 2025
  • A zero-interception season on 1,419 defensive snaps is not a possible outcome

Sample

all 32 defences, 2025 regular season

Method

Count interceptions and fumble recoveries credited to defteam in play-by-play, cross-checked against the player-week feed and against league-wide interceptions thrown.

What we changed

Every NYJ turnover-derived figure is suppressed from cards, briefs and receipts until the feed is refreshed. Nothing on the board depends on it — no projection input keys on takeaways at all — but a number this wrong must not appear anywhere it could be read as real.

Working: docs/receipts/regime-and-possessions-2026.md
Reproduce: count interception==1 grouped by defteam in data/cache/play_by_play_2025.csv