We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
A reasonable thing to believe: if you can shrink the average error in your pitcher projections, your rankings get better. We had a correction that did exactly that. Pitchers were sorted into four groups, high-innings starters, low-innings starters, closers and other relievers, and each group's projected value was multiplied by a single number fitted from past seasons. Tested on seasons the model had not seen, it cut the average miss from 77.36 to 75.32 rank positions, and it won in all five seasons. That is the kind of result you ship.
We held it back anyway, pending four checks. This is the one that killed it.
Nobody drafts off an average error. You draft off a board: an ordered list, where what matters is who sits in your top 30 and whether the guys near the top are actually the guys. A correction can shave two points off the average miss and still hand you a worse board, if the way it shaves them is by overshooting the pitchers at the top.
There is a structural fact here that decides almost everything. The correction multiplies every pitcher in a group by the same number. That cannot reorder anyone inside a group. It can only slide whole groups past each other. So the entire ranking effect boils down to one question: should starters sit higher relative to relievers than they already do? The fitted numbers say yes, by roughly 13 percent. For the held-out 2025 season: high-innings starters 1.301, low-innings starters 1.203, other relievers 1.127, closers 1.091. That single claim is what got tested.
Across five held-out seasons, Spearman correlation, which measures how closely the ranked order matches the realized order, improved in only 2 of 5 seasons and got worse on average, by −0.0030. It got worse in each of the three most recent seasons. Mean absolute rank error also worsened, by +0.202 on average, and the damage grew over time: −0.04, −0.20, +0.04, +0.38, +0.83.
Two numbers looked good and neither holds up. Points captured in the top 30 averaged +69.8, but that average is one season: 2024 at +315, with the others at 0, −3, −47 and +84. Top-20 hits gained 1 to 2 out of 20 in three seasons while top-50 hits netted exactly zero. A gain that vanishes when you widen the window is what noise looks like.
Here is the composition of the predicted top 30 for 2025, by group.
| Name | SP_high | SP_low | closer | RP |
|---|---|---|---|---|
| engine | 15 | 4 | 8 | 3 |
| corrected | 20 | 4 | 4 | 2 |
| actual | 12 | 14 | 2 | 2 |
The engine already crams in too many high-innings starters: 15 where the real answer was 12. The correction pushes that to 20. It moves away from the truth on the group it was already wrong about. It does fix closers, 8 down to 4 against a true 2, but it buys that with a worse starter mix.
And the real miss sits untouched. The realized top 30 held 14 mid-tier starters. The engine found 4. The correction still finds 4, because a single multiplier per group cannot promote mid-tier starters past the high-innings crowd without inflating the high-innings crowd too. One lever, pointed the wrong way.
The numbers above came from one snapshot of the backtest data. That file was regenerated the same day, and rerunning against the newer version gives different figures: engine average miss 80.45, corrected 79.94, a gap of −0.51 with a confidence range from −1.78 to +0.80. That range includes zero, meaning the accuracy gain is no longer distinguishable from nothing. Seasons won drops from 5/5 to 3/5. On rank metrics the fresh picture is mixed rather than a clean fail: Spearman is still 2/5 and flat at +0.0005, but mean absolute rank error improves 4 of 5 seasons and points captured in the top 30 improves in all five, by +240.2. So this is a weaker rejection than first written, and it is recorded that way.
What this does not prove: that the engine's board is fine. It is not. Mid-tier starters are badly under-ranked. It also does not prove that no recalibration can help rankings, only that a flat per-group multiplier provably cannot. Nothing here says anything about hitters or category leagues. And five seasons across 20 slots cannot settle a 1-hit change in top-20 accuracy either way.
A correction that reorders pitchers within groups rather than scaling groups, and that improves Spearman in a majority of held-out seasons and moves top-30 composition toward the real mix, not just the hit count. Both conditions, deliberately, because a single-metric bar is what let this candidate get as far as it did.
Nothing changes in the rankings you see. The correction is closed, not parked. The useful leftover is a named, specific flaw: 4 mid-tier starters in the top 30 where reality held 14. If you are building your own board, that is where to push, and it is a different job entirely.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.