We Thought Our Model Was Paying Ex-Closers for Saves They'll Never Get. We Were Wrong, and the Way We Were Wrong Matters

We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.

A save is not a skill. It is a job title plus a one-run lead in the ninth. A relief win is closer to a coin flip. So the reasonable worry, going into this, was that our projection system quietly treats last year's save total as evidence about next year's save total, and hands a deposed closer a pile of saves he has no route to. That is exactly the kind of thing you want to catch before a draft. We tested it. The wins answer came back clean, the saves answer came back alarming, and then the alarming part turned out to be my own measurement error. All three of those are in here.

Wins: nothing to fix

We built a panel of 3,841 held-out pitcher seasons from 2015 to 2025, pairs of consecutive years with at least 60 outs in both, and asked how much of next year's wins-per-out we could explain. R² is just the share of variation a model captures, where 1.0 would be perfect.

A model built from role alone (workload, share of games started, relief share, team winning percentage) plus skill got 0.0418. Adding the pitcher's own past win rate took it to 0.0460. Own rate and age alone: 0.0144. That is a gain of +0.0041, below the 0.005 threshold we set in advance for calling a term redundant. It is redundant.

Look at the ceiling, though. The best model explains 4.6 percent of the variance. Wins are close to unpredictable. Our starter path already rebuilds wins from context and never looks at a pitcher's own past win rate, so there is nothing to change.

Saves: own rate is genuinely predictive, and genuinely dangerous at the wrong moment

Saves are not wins. Role plus skill gets 0.2218. Add own past save rate and it jumps to 0.5159, a gain of +0.294. Own rate and age alone: 0.4886. Closers usually stay closers, and that term is doing real work.

Now split by what happened to the role. These are biases in saves per 200 outs, positive meaning over-projection:

CohortnROLEROLE+OWNOWN only
was closer, stayed closer133−24.07−9.68−9.64
was closer, lost the role104−2.36+10.53+11.18
was not closer, gained the role105−18.88−18.66−19.99
never closer3,499+1.55+0.61+0.63

On the lost-the-role group, adding own rate swings bias from −2.36 to +10.53, a move of +12.89 against a pre-set harm threshold of +3.0. Harmful. But read the whole table: both simplified models are bad at transitions, and the role-only model under-projects continuing closers by 24 saves per 200 outs. The real problem is forecasting the role, which these stripped-down test models cannot do and our actual engine tries to.

The scary result, and why it was wrong

I then evaluated the live save formula by hand and found it non-monotonic in closer probability: hold everything fixed, raise the odds a pitcher is the closer, and projected saves sometimes went down. Fifteen of 28 grid cells violated it. A pitcher with 30 prior saves who was certain not to be the closer projected 75.5 saves; certain to be the closer, 29.2. At 40 prior saves the non-closer branch hit 100.7, above the MLB single-season record of 62.

The mechanism is real. The formula's "committee" term scales with the pitcher's own prior save volume and with one minus his closer probability, so it reads a former closer's 30 saves as proof he will keep vulturing 30 without the job. A history term also runs at full weight when closer probability is low.

But I fed it inputs it never actually receives. The current-season save rate is the current season, not the prior one, so for a pitcher who lost the role it already reflects the loss. It is also stabilized toward a very low baseline; my probe used a raw rate roughly 65 times that target. That single error produced essentially all the explosive numbers.

What the real board says

Checking the shipped projections for 672 pitchers, controlling for workload by taking relievers projected for 45 to 85 innings:

Closer probabilitynmean prior SVmean projected IPmean projected SV
under 0.301220.461.40.4
0.30 to 0.60851.560.83.3
0.85 or higher4012.862.120.1

Cleanly increasing. No flatness, no inversion. Edwin Díaz, 28 prior saves and a 0.20 closer probability, projects 4.0 saves. Robert Suarez, 40 prior saves, 0.49 closer probability, 48.4 innings, projects 6.5. Board-wide maximum is 47.6, with zero pitchers above the single-season record.

A second claim of mine also failed: I said the leverage input saturates at its maximum for anyone with five or more prior saves. On the live board only 3.4 percent of pitchers sit at the cap, and even among pitchers with 15-plus prior saves only 44 percent do, mean 0.656.

What this does and does not establish

The wins finding stands; it never depended on the probe. The formula genuinely is non-monotonic over the range I tested, but that range is unreachable with the inputs it actually gets. That makes it a latent fragility, not a live defect: it would matter only if the stabilization or the input meaning changed. The panel harm result still describes the simplified test models honestly, but its bearing on the shipped engine is now weak, because the engine does not use own rate the way those models did.

No changes are being made. A cheap structural test asserting that projected saves never fall as closer probability rises is worth adding as insurance. Nothing else.

One loose thread, recorded rather than asserted: projected saves across the board total 1,472 against an MLB reality near 1,200. Some over-allocation is structural, since two teammates can each carry partial closer probability, so this needs team-level normalization before it means anything.

The lesson

Evaluating a formula directly feels like the strongest possible evidence. It is deterministic, nothing is estimated. That is exactly what makes it dangerous when you assume what the inputs mean instead of tracing them. The check that overturned this cost one script and read a file we already publish, and it should have run before the writeup, not after.

All articles · Redraft rankings · Dynasty rankings

Skip to main content

Redraft Rest of Season

Sorted by -
Loading history…
Official MiLB Prospect Rankings

Official MiLB Prospect Rankings

Loading…
Loading rankings…
Overview

Operations Dashboard

Portfolio and league analytics.
Workspace ready
Use Sync to import a roster.

2026 FYPD Class

Players entering the MiLB system for the first time: the 2026 MLB Draft class and first-time international signees. Ranked by dynasty value: a measured expectation from draft slot or debut production, adjusted by the age curve. For established prospects, see the Prospect board.

Prospect Player Rankings

Top-100 + all 30 org Top-30 lists · ranked by projected Prospect Value · AAA Statcast tools →
Loading prospect rankings…
Loading study…

Methodology

Every signal, formula, and data source the model uses for player evaluation - organized by product.

Loading methodology…

Risers & Fallers - last 30 days