Your Projections Trust Batting Average Four Times More Than They Should

We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.

You could reasonably have believed batting average was the one traditional stat fantasy had already made peace with. The Hidden Game of Baseball spent its opening arguing that AVG is a poor measure of a hitter, that RBI counts opportunity, that runs are not individual, that pitcher wins mean almost nothing. Fine. But roto scores AVG as a category, so agreeing it measures little does not excuse you from forecasting it well. That gap is worth testing, and testing it turned up something that had nothing to do with the philosophy and everything to do with one number buried in the projection.

The book's thesis, measured

Predicting next-year AVG from a single input, on 693 hitters with at least 250 plate appearances in both seasons, holding out one season at a time:

predictorout-of-sample MAE (average miss)R² (share of variation explained)
past AVG (outcome)0.019950.1719
past xBA (skill)0.019620.2195
past BABIP (luck)0.021720.0275
AVG + full skill set0.019350.2532

A skill measure predicts next year's batting average better than batting average does: R² 0.2195 against 0.1719, a 28% relative gain. Luck-driven BABIP carries almost nothing. That is the thesis, shown rather than asserted.

The obvious fix failed its own test

So add xBA to the projection? The projection does not carry raw past AVG anyway. It uses a Marcel: recency-weighted AVG pulled toward the league mean by a fixed amount of imaginary "ghost" at-bats, then age-adjusted, with no Statcast input at all. Against that stronger baseline, on 465 hitters across two folds:

modelout-of-sample MAEvs Marcel
Marcel, outcome only0.02028—
Marcel + xBA0.019822.3%
Marcel + xBA + K% + speed0.019812.3%

The bar was set in advance at 3% MAE improvement, because Marcel's own edge over a naive carry is about 11%. 2.3% fails it. The signal is real and pointing the right way. It does not clear the gate on a thin sample, and that is the honest result.

The real finding: how hard AVG gets regressed

The regression amount was set at 700 ghost at-bats and labeled validated. It had been tested against 300, 500, 700, 900 and 1200, and improved all the way to the largest value tried. An optimum sitting on the edge of your grid is a fact about the grid, not a result.

Widen it and the true optimum is near 3000, confirmed twice independently:

ghost ABMAE (in-house)MAE (independent)
700 (shipped)0.02160.02159
12000.02110.02105
26000.02070.02068
30000.02070.02068
45000.02080.02079
60000.02100.02101

The shipped setting was 4.2 to 4.4% worse than optimum, on 804 pairs over three target seasons. In plain terms: an individual's batting average deserves roughly four times less trust than the projection was giving it, because most of it is noise.

Why the fix stopped short of the optimum

There is a cost. Heavier regression squashes everyone toward the league mean, which helps absolute error and hurts ordering. Correlation falls from 0.466 at 700 to 0.458 at 3000. Roto cares about both: who helps your average, and by how much.

ghost ABMAErMAE gainspread kept
700 (was)0.021630.466—100%
**1200 (shipped)**0.021130.4662.3%88%
16000.020920.4653.3%81%
22000.020780.4623.9%72%
3000 (MAE optimum)0.020730.4584.2%63%

1200 is the last point where correlation is provably unchanged: 55% of the available MAE gain at zero ranking cost. At 3000 the compression is concrete and indefensible for a category, with Arraez falling from .303 to .289 and Gallo rising from .188 to .214, a 37% flattening of exactly the ordering the category consumes. Projecting the best pure contact hitter in baseball at .289 to buy 1.9% more MAE is the metric choosing against the product.

The same defect hit home runs, and that one reaches points leagues

This document originally claimed home run rate was fine, with an interior optimum near 350 and 350 shipped. That was wrong. It was read off a table ending at 450, where 350 and 450 tied to four decimals. The HR grid had the identical edge problem.

constantshippedoptimumshipped isr, shipped → optimum
HR ghost AB35011002.0% worse0.623 → 0.623
AVG ghost AB70030004.3% worse0.466 → 0.458

HR is the cleaner win despite the smaller number: correlation is identical, so there is no trade-off column at all. MAE moves 0.00944 to 0.00926. Both were moved: HR to 1100, AVG to 1200.

The HR change matters more than it looks, because the home run line feeds a total-bases consistency floor, and that floor writes projected points. Regenerating both ways across 774 players:

groupchange in HR per 600 PA
Judge−2.7
Stanton−3.3
Montgomery−3.2
top-10 power, mean−2.41
all 471 established batters, mean+0.29 (range −3.5 to +3.8)

Elite power comes down, the middle comes up. The floor is roughly four times HR, so it weakens by about 11 total bases for the elite power group, the group the floor exists to protect. Small against a floor of about 277, and the 2.0% accuracy gain applies to every batter. Still, that is the group to check. The AVG change is narrower: it touches the categories line and a display column, nothing that reaches projected points.

What this does not prove

It does not prove xBA is worthless, only that it misses a 3% bar on 465 hitters. A bigger panel could clear it, and the skill data only exists for four seasons.

It does not prove 3000 is right for the live projection. These numbers come from a faithful re-implementation and an offline test bed. The live version also carries an age adjustment and a different history window.

And the HR change was never graded against realised projected points on a live board, only against home run rate accuracy offline. It shipped on the backtest plus the measured side effect, not on a full live check.

What would change the answer

For AVG: a live run showing 700 matches or beats 3000 on held-out seasons once the age adjustment and history window are in play. If the offline test and the live pipeline disagree about where the optimum sits, the finding is refuted.

For HR: a live board comparison showing elite power bats losing materially more projected points than the roughly 11 total bases the floor arithmetic predicts. Reverting is one number and a rebuild.

Nothing here says anything about pitchers, and nothing here grades how AVG is scored downstream. Extending AVG to 1600, 2200 or 3000 remains available if the ranking cost is judged acceptable. The curve above is the whole decision.

All articles · Redraft rankings · Dynasty rankings

Skip to main content

Redraft Rest of Season

Sorted by -
Loading history…
Official MiLB Prospect Rankings

Official MiLB Prospect Rankings

Loading…
Loading rankings…
Overview

Operations Dashboard

Portfolio and league analytics.
Workspace ready
Use Sync to import a roster.

2026 FYPD Class

Players entering the MiLB system for the first time: the 2026 MLB Draft class and first-time international signees. Ranked by dynasty value: a measured expectation from draft slot or debut production, adjusted by the age curve. For established prospects, see the Prospect board.

Prospect Player Rankings

Top-100 + all 30 org Top-30 lists · ranked by projected Prospect Value · AAA Statcast tools →
Loading prospect rankings…
Loading study…

Methodology

Every signal, formula, and data source the model uses for player evaluation - organized by product.

Loading methodology…

Risers & Fallers - last 30 days