We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
A save is not a skill. It is a job title plus a one-run lead in the ninth. A relief win is closer to a coin flip. So the reasonable worry, going into this, was that our projection system quietly treats last year's save total as evidence about next year's save total, and hands a deposed closer a pile of saves he has no route to. That is exactly the kind of thing you want to catch before a draft. We tested it. The wins answer came back clean, the saves answer came back alarming, and then the alarming part turned out to be my own measurement error. All three of those are in here.
We built a panel of 3,841 held-out pitcher seasons from 2015 to 2025, pairs of consecutive years with at least 60 outs in both, and asked how much of next year's wins-per-out we could explain. R² is just the share of variation a model captures, where 1.0 would be perfect.
A model built from role alone (workload, share of games started, relief share, team winning percentage) plus skill got 0.0418. Adding the pitcher's own past win rate took it to 0.0460. Own rate and age alone: 0.0144. That is a gain of +0.0041, below the 0.005 threshold we set in advance for calling a term redundant. It is redundant.
Look at the ceiling, though. The best model explains 4.6 percent of the variance. Wins are close to unpredictable. Our starter path already rebuilds wins from context and never looks at a pitcher's own past win rate, so there is nothing to change.
Saves are not wins. Role plus skill gets 0.2218. Add own past save rate and it jumps to 0.5159, a gain of +0.294. Own rate and age alone: 0.4886. Closers usually stay closers, and that term is doing real work.
Now split by what happened to the role. These are biases in saves per 200 outs, positive meaning over-projection:
| Cohort | n | ROLE | ROLE+OWN | OWN only |
|---|---|---|---|---|
| was closer, stayed closer | 133 | −24.07 | −9.68 | −9.64 |
| was closer, lost the role | 104 | −2.36 | +10.53 | +11.18 |
| was not closer, gained the role | 105 | −18.88 | −18.66 | −19.99 |
| never closer | 3,499 | +1.55 | +0.61 | +0.63 |
On the lost-the-role group, adding own rate swings bias from −2.36 to +10.53, a move of +12.89 against a pre-set harm threshold of +3.0. Harmful. But read the whole table: both simplified models are bad at transitions, and the role-only model under-projects continuing closers by 24 saves per 200 outs. The real problem is forecasting the role, which these stripped-down test models cannot do and our actual engine tries to.
I then evaluated the live save formula by hand and found it non-monotonic in closer probability: hold everything fixed, raise the odds a pitcher is the closer, and projected saves sometimes went down. Fifteen of 28 grid cells violated it. A pitcher with 30 prior saves who was certain not to be the closer projected 75.5 saves; certain to be the closer, 29.2. At 40 prior saves the non-closer branch hit 100.7, above the MLB single-season record of 62.
The mechanism is real. The formula's "committee" term scales with the pitcher's own prior save volume and with one minus his closer probability, so it reads a former closer's 30 saves as proof he will keep vulturing 30 without the job. A history term also runs at full weight when closer probability is low.
But I fed it inputs it never actually receives. The current-season save rate is the current season, not the prior one, so for a pitcher who lost the role it already reflects the loss. It is also stabilized toward a very low baseline; my probe used a raw rate roughly 65 times that target. That single error produced essentially all the explosive numbers.
Checking the shipped projections for 672 pitchers, controlling for workload by taking relievers projected for 45 to 85 innings:
| Closer probability | n | mean prior SV | mean projected IP | mean projected SV |
|---|---|---|---|---|
| under 0.30 | 122 | 0.4 | 61.4 | 0.4 |
| 0.30 to 0.60 | 85 | 1.5 | 60.8 | 3.3 |
| 0.85 or higher | 40 | 12.8 | 62.1 | 20.1 |
Cleanly increasing. No flatness, no inversion. Edwin Díaz, 28 prior saves and a 0.20 closer probability, projects 4.0 saves. Robert Suarez, 40 prior saves, 0.49 closer probability, 48.4 innings, projects 6.5. Board-wide maximum is 47.6, with zero pitchers above the single-season record.
A second claim of mine also failed: I said the leverage input saturates at its maximum for anyone with five or more prior saves. On the live board only 3.4 percent of pitchers sit at the cap, and even among pitchers with 15-plus prior saves only 44 percent do, mean 0.656.
The wins finding stands; it never depended on the probe. The formula genuinely is non-monotonic over the range I tested, but that range is unreachable with the inputs it actually gets. That makes it a latent fragility, not a live defect: it would matter only if the stabilization or the input meaning changed. The panel harm result still describes the simplified test models honestly, but its bearing on the shipped engine is now weak, because the engine does not use own rate the way those models did.
No changes are being made. A cheap structural test asserting that projected saves never fall as closer probability rises is worth adding as insurance. Nothing else.
One loose thread, recorded rather than asserted: projected saves across the board total 1,472 against an MLB reality near 1,200. Some over-allocation is structural, since two teammates can each carry partial closer probability, so this needs team-level normalization before it means anything.
Evaluating a formula directly feels like the strongest possible evidence. It is deterministic, nothing is estimated. That is exactly what makes it dangerous when you assume what the inputs mean instead of tracing them. The check that overturned this cost one script and read a file we already publish, and it should have run before the writeup, not after.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.