We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
A reasonable belief: if a projection system nudges players toward a fixed reference point, then any rule that decides who gets spared that nudge ought to use the same reference point. Ours does not. The regression step pulls every player toward a career-plate-appearance anchor of 160, meaning it drags players above 160 down and pushes players below 160 up. The rule that exempts certain young hitters from that pull, however, tests a threshold of 100. So a player sitting between 100 and 160 is being "spared" a correction that would have moved him up the board. That is backwards, and nobody on our side disputes it. On the live board it affects 10 rows out of 1,322.
We wrote the test down before we ran it, including the minimum sample we would accept: at least 50 comparison pairs, since anything smaller cannot separate a real effect from noise. Over 2021 to 2025 the replay produced 1,705 player-season pairs, of which 60 cleared the pedigree, age and career-plate-appearance conditions, and only 8 landed in the band where a boundary of 100 and a boundary of 160 actually disagree. The average size of a projection miss, regardless of direction, came in at 87.04 with the boundary at 100 and 87.68 with it at 160.
Eight pairs. The pre-registered rule at n=8 is "ship nothing," and the current boundary was nominally the better of the two anyway.
Stretching back to 2016 through 2025, with 2020 excluded, was not the test we registered, so treat it as exploratory. It gave 3,223 pairs, 140 clearing the gate conditions, and 17 in the disagreement band. Now the error runs 81.87 at a boundary of 100 against 79.73 at 160, with 6 of 9 seasons favoring 160. Doubling the window moved us from 8 pairs to 17 and reversed the conclusion outright.
That is the strongest case against shipping, and it is also the strongest case for the fix, depending on which table you trust. I read two underpowered tests pointing in opposite directions as the signature of noise, not of a real effect waiting for more data. If one more season of history can flip the sign, the sign is not information.
The gate is structurally rare. Only 140 of 3,223 historical pairs, 4.3%, clear pedigree plus age plus career plate appearances, and only 17 sit where the two boundaries disagree. The live board shows the same scarcity at 10 rows of 1,322, so this is not a quirk of the years we chose. Any boundary change here is evidence-free by construction.
Two further limits surfaced from reading the logic rather than from the numbers. The gate drives two separate lifts, one tied to a 900 career-plate-appearance figure and one to 800, so moving the boundary moves both. And the second of those lifts is wired specifically to 2026, while the replay only ever feeds it seasons earlier than the one being projected. It can never fire in a historical test. That is a permanent blind spot, not a bug in the run.
One trap for anyone repeating this: the pedigree lookup keyed on player name returns today's grades with no season attached, which would apply 2026 pedigree to a 2021 decision. We used the year-aware lookup instead.
Nothing shipped. The rule still tests 100. We recorded the measurement, the two limits and the trap alongside it, with an instruction not to re-run this expecting a cleaner answer without a better instrument.
What the evidence does not establish: that 100 is correct, and it almost certainly is not, since the coherence argument stands unrebutted. That 160 is better, because the wider window says so at n=17 and the registered window says the opposite at n=8. And anything at all about the proven-debut lift, which the replay cannot reach.
Shipping on coherence alone was available. We declined it. A change that cannot be shown to help is not obviously worth 10 rows of movement.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.