The Fix That Made Our Projections More Accurate and Our Rankings Worse

We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.

A reasonable thing to believe: if you can shrink the average error in your pitcher projections, your rankings get better. We had a correction that did exactly that. Pitchers were sorted into four groups, high-innings starters, low-innings starters, closers and other relievers, and each group's projected value was multiplied by a single number fitted from past seasons. Tested on seasons the model had not seen, it cut the average miss from 77.36 to 75.32 rank positions, and it won in all five seasons. That is the kind of result you ship.

We held it back anyway, pending four checks. This is the one that killed it.

Why average error is the wrong scoreboard

Nobody drafts off an average error. You draft off a board: an ordered list, where what matters is who sits in your top 30 and whether the guys near the top are actually the guys. A correction can shave two points off the average miss and still hand you a worse board, if the way it shaves them is by overshooting the pitchers at the top.

There is a structural fact here that decides almost everything. The correction multiplies every pitcher in a group by the same number. That cannot reorder anyone inside a group. It can only slide whole groups past each other. So the entire ranking effect boils down to one question: should starters sit higher relative to relievers than they already do? The fitted numbers say yes, by roughly 13 percent. For the held-out 2025 season: high-innings starters 1.301, low-innings starters 1.203, other relievers 1.127, closers 1.091. That single claim is what got tested.

It failed on the metrics that matter

Across five held-out seasons, Spearman correlation, which measures how closely the ranked order matches the realized order, improved in only 2 of 5 seasons and got worse on average, by −0.0030. It got worse in each of the three most recent seasons. Mean absolute rank error also worsened, by +0.202 on average, and the damage grew over time: −0.04, −0.20, +0.04, +0.38, +0.83.

Two numbers looked good and neither holds up. Points captured in the top 30 averaged +69.8, but that average is one season: 2024 at +315, with the others at 0, −3, −47 and +84. Top-20 hits gained 1 to 2 out of 20 in three seasons while top-50 hits netted exactly zero. A gain that vanishes when you widen the window is what noise looks like.

The mechanism, which is the actual story

Here is the composition of the predicted top 30 for 2025, by group.

NameSP_highSP_lowcloserRP
engine15483
corrected20442
actual121422

The engine already crams in too many high-innings starters: 15 where the real answer was 12. The correction pushes that to 20. It moves away from the truth on the group it was already wrong about. It does fix closers, 8 down to 4 against a true 2, but it buys that with a worse starter mix.

And the real miss sits untouched. The realized top 30 held 14 mid-tier starters. The engine found 4. The correction still finds 4, because a single multiplier per group cannot promote mid-tier starters past the high-innings crowd without inflating the high-innings crowd too. One lever, pointed the wrong way.

The honest caveats

The numbers above came from one snapshot of the backtest data. That file was regenerated the same day, and rerunning against the newer version gives different figures: engine average miss 80.45, corrected 79.94, a gap of −0.51 with a confidence range from −1.78 to +0.80. That range includes zero, meaning the accuracy gain is no longer distinguishable from nothing. Seasons won drops from 5/5 to 3/5. On rank metrics the fresh picture is mixed rather than a clean fail: Spearman is still 2/5 and flat at +0.0005, but mean absolute rank error improves 4 of 5 seasons and points captured in the top 30 improves in all five, by +240.2. So this is a weaker rejection than first written, and it is recorded that way.

What this does not prove: that the engine's board is fine. It is not. Mid-tier starters are badly under-ranked. It also does not prove that no recalibration can help rankings, only that a flat per-group multiplier provably cannot. Nothing here says anything about hitters or category leagues. And five seasons across 20 slots cannot settle a 1-hit change in top-20 accuracy either way.

What would change the answer

A correction that reorders pitchers within groups rather than scaling groups, and that improves Spearman in a majority of held-out seasons and moves top-30 composition toward the real mix, not just the hit count. Both conditions, deliberately, because a single-metric bar is what let this candidate get as far as it did.

What you do with this

Nothing changes in the rankings you see. The correction is closed, not parked. The useful leftover is a named, specific flaw: 4 mid-tier starters in the top 30 where reality held 14. If you are building your own board, that is where to push, and it is a different job entirely.

All articles · Redraft rankings · Dynasty rankings

Skip to main content

Redraft Rest of Season

Sorted by -
Loading history…
Official MiLB Prospect Rankings

Official MiLB Prospect Rankings

Loading…
Loading rankings…
Overview

Operations Dashboard

Portfolio and league analytics.
Workspace ready
Use Sync to import a roster.

2026 FYPD Class

Players entering the MiLB system for the first time: the 2026 MLB Draft class and first-time international signees. Ranked by dynasty value: a measured expectation from draft slot or debut production, adjusted by the age curve. For established prospects, see the Prospect board.

Prospect Player Rankings

Top-100 + all 30 org Top-30 lists · ranked by projected Prospect Value · AAA Statcast tools →
Loading prospect rankings…
Loading study…

Methodology

Every signal, formula, and data source the model uses for player evaluation - organized by product.

Loading methodology…

Risers & Fallers - last 30 days