We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
You could reasonably have believed batting average was the one traditional stat fantasy had already made peace with. The Hidden Game of Baseball spent its opening arguing that AVG is a poor measure of a hitter, that RBI counts opportunity, that runs are not individual, that pitcher wins mean almost nothing. Fine. But roto scores AVG as a category, so agreeing it measures little does not excuse you from forecasting it well. That gap is worth testing, and testing it turned up something that had nothing to do with the philosophy and everything to do with one number buried in the projection.
Predicting next-year AVG from a single input, on 693 hitters with at least 250 plate appearances in both seasons, holding out one season at a time:
| predictor | out-of-sample MAE (average miss) | R² (share of variation explained) |
|---|---|---|
| past AVG (outcome) | 0.01995 | 0.1719 |
| past xBA (skill) | 0.01962 | 0.2195 |
| past BABIP (luck) | 0.02172 | 0.0275 |
| AVG + full skill set | 0.01935 | 0.2532 |
A skill measure predicts next year's batting average better than batting average does: R² 0.2195 against 0.1719, a 28% relative gain. Luck-driven BABIP carries almost nothing. That is the thesis, shown rather than asserted.
So add xBA to the projection? The projection does not carry raw past AVG anyway. It uses a Marcel: recency-weighted AVG pulled toward the league mean by a fixed amount of imaginary "ghost" at-bats, then age-adjusted, with no Statcast input at all. Against that stronger baseline, on 465 hitters across two folds:
| model | out-of-sample MAE | vs Marcel |
|---|---|---|
| Marcel, outcome only | 0.02028 | — |
| Marcel + xBA | 0.01982 | 2.3% |
| Marcel + xBA + K% + speed | 0.01981 | 2.3% |
The bar was set in advance at 3% MAE improvement, because Marcel's own edge over a naive carry is about 11%. 2.3% fails it. The signal is real and pointing the right way. It does not clear the gate on a thin sample, and that is the honest result.
The regression amount was set at 700 ghost at-bats and labeled validated. It had been tested against 300, 500, 700, 900 and 1200, and improved all the way to the largest value tried. An optimum sitting on the edge of your grid is a fact about the grid, not a result.
Widen it and the true optimum is near 3000, confirmed twice independently:
| ghost AB | MAE (in-house) | MAE (independent) |
|---|---|---|
| 700 (shipped) | 0.0216 | 0.02159 |
| 1200 | 0.0211 | 0.02105 |
| 2600 | 0.0207 | 0.02068 |
| 3000 | 0.0207 | 0.02068 |
| 4500 | 0.0208 | 0.02079 |
| 6000 | 0.0210 | 0.02101 |
The shipped setting was 4.2 to 4.4% worse than optimum, on 804 pairs over three target seasons. In plain terms: an individual's batting average deserves roughly four times less trust than the projection was giving it, because most of it is noise.
There is a cost. Heavier regression squashes everyone toward the league mean, which helps absolute error and hurts ordering. Correlation falls from 0.466 at 700 to 0.458 at 3000. Roto cares about both: who helps your average, and by how much.
| ghost AB | MAE | r | MAE gain | spread kept |
|---|---|---|---|---|
| 700 (was) | 0.02163 | 0.466 | — | 100% |
| **1200 (shipped)** | 0.02113 | 0.466 | 2.3% | 88% |
| 1600 | 0.02092 | 0.465 | 3.3% | 81% |
| 2200 | 0.02078 | 0.462 | 3.9% | 72% |
| 3000 (MAE optimum) | 0.02073 | 0.458 | 4.2% | 63% |
1200 is the last point where correlation is provably unchanged: 55% of the available MAE gain at zero ranking cost. At 3000 the compression is concrete and indefensible for a category, with Arraez falling from .303 to .289 and Gallo rising from .188 to .214, a 37% flattening of exactly the ordering the category consumes. Projecting the best pure contact hitter in baseball at .289 to buy 1.9% more MAE is the metric choosing against the product.
This document originally claimed home run rate was fine, with an interior optimum near 350 and 350 shipped. That was wrong. It was read off a table ending at 450, where 350 and 450 tied to four decimals. The HR grid had the identical edge problem.
| constant | shipped | optimum | shipped is | r, shipped → optimum |
|---|---|---|---|---|
| HR ghost AB | 350 | 1100 | 2.0% worse | 0.623 → 0.623 |
| AVG ghost AB | 700 | 3000 | 4.3% worse | 0.466 → 0.458 |
HR is the cleaner win despite the smaller number: correlation is identical, so there is no trade-off column at all. MAE moves 0.00944 to 0.00926. Both were moved: HR to 1100, AVG to 1200.
The HR change matters more than it looks, because the home run line feeds a total-bases consistency floor, and that floor writes projected points. Regenerating both ways across 774 players:
| group | change in HR per 600 PA |
|---|---|
| Judge | −2.7 |
| Stanton | −3.3 |
| Montgomery | −3.2 |
| top-10 power, mean | −2.41 |
| all 471 established batters, mean | +0.29 (range −3.5 to +3.8) |
Elite power comes down, the middle comes up. The floor is roughly four times HR, so it weakens by about 11 total bases for the elite power group, the group the floor exists to protect. Small against a floor of about 277, and the 2.0% accuracy gain applies to every batter. Still, that is the group to check. The AVG change is narrower: it touches the categories line and a display column, nothing that reaches projected points.
It does not prove xBA is worthless, only that it misses a 3% bar on 465 hitters. A bigger panel could clear it, and the skill data only exists for four seasons.
It does not prove 3000 is right for the live projection. These numbers come from a faithful re-implementation and an offline test bed. The live version also carries an age adjustment and a different history window.
And the HR change was never graded against realised projected points on a live board, only against home run rate accuracy offline. It shipped on the backtest plus the measured side effect, not on a full live check.
For AVG: a live run showing 700 matches or beats 3000 on held-out seasons once the age adjustment and history window are in play. If the offline test and the live pipeline disagree about where the optimum sits, the finding is refuted.
For HR: a live board comparison showing elite power bats losing materially more projected points than the roughly 11 total bases the floor arithmetic predicts. Reverting is one number and a rebuild.
Nothing here says anything about pitchers, and nothing here grades how AVG is scored downstream. Extending AVG to 1600, 2200 or 3000 remains available if the ranking cost is judged acceptable. The curve above is the whole decision.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.