Two Corrections Fighting Each Other: What We Found Inside Our Pitcher ERA-Luck Math

We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.

There is a claim you have probably heard: reliever ERA is noise. The version that set this off put numbers on it, saying reliever ERA carries essentially no year-to-year signal (R² of 0.003, where R² is the share of next year's variation explained by this year's) while SIERA carries 0.13. Value relievers on process, not results. It is a reasonable thing to believe, and our projection system already leans that way. What it had never checked is whether one setting, tuned mostly on starters, was quietly under-correcting relievers. That was the question. The answer turned into something more interesting.

The reliever asymmetry is real

Building pairs of consecutive seasons from 2015 to 2025, requiring 120+ outs (40+ innings) in both years, gives 2,321 pairs: 1,206 starters, 745 relievers, 370 swing arms, with role assigned from the first year so nothing uses information a projection would not have.

bucketERA → ERA R²xERA → ERA R²xERA → xERA R²
SP (n=1,206)0.03960.08530.1571
RP (n=745)0.00540.03370.0618
swing (n=370)0.01360.04700.0845

Reliever ERA is about 7× less self-predictive than starter ERA, and 0.0054 sits close to the outside claim's 0.003. That part replicates cleanly.

The headline 40× estimator advantage does not. That figure compares ERA predicting ERA against SIERA predicting SIERA, which is a same-metric comparison and says nothing about forecasting outcomes. The comparable ratio here is about 11×. And the number that actually matters for a projection, the estimator against next year's real ERA, is 0.0337 for relievers. Better than ERA. Still close to useless in absolute terms. The defensible sentence is: reliever ERA is not well predicted by anything in this data. Role churn is not the culprit either, since 87.2% of relievers who survive into the panel are still relievers the next year.

The isolated test said to turn the knob down for relievers

Our system nudges projected earned runs toward or away from observed ERA using a strength setting shipped at 10%. Testing that formula on its own, and comparing against the shipped 10% with a paired bootstrap (resampling the same pitchers thousands of times to get an error range), plus leave-one-season-out testing where the strength is picked on other seasons and scored on the held-out one:

bucket0% vs 10%LOSO pickLOSO Δ MAE
POOLED−0.0048 [−0.0110, +0.0017] ns0% in 8/10−0.0036
SP+0.0012 [−0.0085, +0.0108] ns5% in 10/10−0.0021
RP−0.0099 [−0.0192, −0.0012] SIG0% in 8/8−0.0103
swing−0.0141 [−0.0311, +0.0036] ns0% in 8/8−0.0169

Negative means the candidate beat shipped. The pooled verdict of "not significant" is a null starter effect and a significant reliever effect averaging into mush. Best strength: 10% for starters, 0% for relievers. The effect is small: for relievers it moves earned-run MAE from 6.709 to 6.652, about 0.06 earned runs over a season, with rank correlation improving from 0.1837 to 0.1976.

Then we looked at whether that knob acts alone. It does not.

Before recommending a change, we checked whether anything else in the pipeline was already working on the same signal. Three separate corrections exist, not one. L1 multiplies projected earned runs toward actual ERA. L2 blends projected ERA 70/20 toward current-year xERA. L3 adjusts projected points toward xFIP, falling back to SIERA. Only L1 had ever been tested.

On a settled board of 683 pitchers, L1 fires on 661 (96.8%), L2 on 97 (14.2%), L3 on 430 (63.0%, xFIP every time, the SIERA fallback never fires). L1 and L3 fire together on 425 pitchers, 62.2%. All three on 81, 11.9%.

Then the split by whether a pitcher's ERA beat his xERA:

Namelucky (ERA < xERA)unlucky (ERA > xERA)
L1+2.89 pts (n=360)−3.23 pts (n=301)
L3−4.07 pts (n=234)+4.37 pts (n=196)

They are exact opposites. L1 rewards the pitcher who got lucky; L3 punishes him. On the 425 co-firing pitchers the signs oppose in 352 cases, 82.8%, and 62.6% of the combined magnitude cancels: mean |L1| of 3.25 plus mean |L3| of 4.22 gives 7.47 points of gross adjustment collapsing to a mean absolute net of 2.80. Mean net is −0.02 points, median −0.52.

What survives is close to a coin flip: 221 pitchers (52.0%) end up moved in the estimator direction, 181 (42.6%) in the opposite direction, 23 (5.4%) under a quarter point. The largest surviving distortions, where the cancellation happened to fail, are Garrett Crochet at net +10.3, Cal Quantrill +8.7 and Hunter Greene +8.2.

One thing here is objectively wrong and now fixed: the note attached to L1 described it as buying the unlucky pitcher and selling the lucky one. The code does the opposite. L3 is what implements that stated intent. Whether L1's behavior is wrong is genuinely arguable, since the base it multiplies is skill-derived and pulling a skill projection partway toward results is a defensible idea under a different name. The description was still backwards, and someone tuning it from that description would have moved it the wrong way.

The test that would have made this matter, and how it failed

We committed in advance to the thing that would settle it: turn the layers off in the full pipeline and see whether the board or the realized accuracy actually moves. If it does not, this is cosmetic, not a defect.

It does not. Accuracy, on 1,764 pairs with a fresh rebuild for each of the 8 on/off combinations:

variantMAEΔ vs shipped95% CIPearsonSpearman
shipped77.747——.4675.4133
L1 off77.646−0.102[−0.289, +0.076] ns.4704.4150
L3 off77.927+0.180[−0.010, +0.382] ns.4621.4078
L1+L3 off77.823+0.076[−0.165, +0.321] ns.4639.4107

Every interval includes zero. The direction is consistent across all three metrics: removing L1 helps slightly, removing L3 hurts slightly. Both at roughly 0.2% of MAE. The idea survives in sign and dies in size.

Ranks were re-run after we found and froze a source of non-reproducibility in the board loading path, using 2 loads per variant, all 5 variants byte-identical:

variantmean abs Δrankmax≥1≥10≥25Spearman vs shipped
L1 off0.769.4147000.999625
L2 off0.5822.778900.999583
L3 off0.3011.247100.999868
L1+L3 off0.929.1179000.999522

No pitcher moves 25 ranks under any variant. Mean displacement is under one rank. The earlier claim of "Spearman ≥ 0.9996" does not reproduce; the correct bound is ≥ 0.9995, and we state the corrected figure rather than defend the old one. The withdrawn numbers were essentially right: the largest change in mean absolute rank movement across variants was 0.039 ranks.

L2 is the exception. Re-validated on the current pipeline, turning it off costs +0.1925 in ERA-MAE, with a range of [+0.1375, +0.2489]. L2 earns its place and should not be touched.

What this does not prove

The reliever finding does not survive the full pipeline. Killing L1 helps starters more (−0.263 on a 101.25 MAE base) than relievers (−0.031 on 67.35), in both absolute and relative terms, the reverse of the isolated result. Neither is significant. The honest reading is that the reliever optimum was a product of testing one pathway alone, and it is withdrawn as actionable. The raw persistence asymmetry stands as a fact about pitchers; the tuning recommendation does not.

The reliever panel requires 40+ innings in both seasons. A reliever who loses his job and throws 30 innings drops out entirely, so the sample skews toward role-stable arms and, if anything, understates how unpredictable real relievers are. The co-firing measurements come from a snapshot of today's player pool, not a backtest; signs are exact, magnitudes approximate, because later normalization steps can rescale them.

We also had to withdraw a number. We initially reported that killing L1 moved final board points by 0.45 against a gross adjustment of 3.25, roughly 7× attenuation. That was measured by differencing cold loads, and the cold-load path turned out to be non-deterministic, with drift rather than random jitter across identical runs. The 7× is unmeasured, not merely imprecise. The lesson behind it, that a correction's size where it is applied is not its size on the board you actually read, is plausible and unestablished.

Still open: which direction the net ERA-luck treatment should point, whether the 52/43 residual split is signal or noise, whether L2 interacts with the other two, and whether the reliever result holds without the innings gate.

What changes for you

Nothing. No production change is warranted. L1 and L3 have no accuracy or rank case for removal; L2 is actively validated and its shipped shares are optimal. The description error is corrected. We did not ship a role-conditional branch for relievers: it buys roughly 0.004 MAE over the simpler global setting, which does not clear the bar for adding complexity.

The interim suggestion to drop L1's strength from 0.10 to 0.05 is retracted. It was measured in isolation, and live, most of it gets cancelled downstream. If anyone wants to revisit this, the order is: decide the intended direction first, test the surviving layer in the full pipeline, re-run the starter/reliever split against the composed system rather than one pathway, and only then touch a strength.

All articles · Redraft rankings · Dynasty rankings

Skip to main content

Redraft Rest of Season

Sorted by -
Loading history…
Official MiLB Prospect Rankings

Official MiLB Prospect Rankings

Loading…
Loading rankings…
Overview

Operations Dashboard

Portfolio and league analytics.
Workspace ready
Use Sync to import a roster.

2026 FYPD Class

Players entering the MiLB system for the first time: the 2026 MLB Draft class and first-time international signees. Ranked by dynasty value: a measured expectation from draft slot or debut production, adjusted by the age curve. For established prospects, see the Prospect board.

Prospect Player Rankings

Top-100 + all 30 org Top-30 lists · ranked by projected Prospect Value · AAA Statcast tools →
Loading prospect rankings…
Loading study…

Methodology

Every signal, formula, and data source the model uses for player evaluation - organized by product.

Loading methodology…

Risers & Fallers - last 30 days