We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
Batting order matters. Everyone knows it: a hitter moved from seventh to second gains runs, and one moved out of the middle of the order loses RBI. So a projection system that guesses a hitter's lineup slot from his own projected rates, which is circular reasoning, looks like an obvious thing to fix. Feed it the slot he actually hit in, and the runs and RBI numbers should get better. That was the claim, and it had never been tested on its own.
We tested it. It failed.
We compared four versions on 919 season-to-season player pairs across four seasons the model had never seen, using both mean absolute error (MAE, the average miss in runs or RBI) and rank correlation, or ρ, which measures whether the players are ordered correctly and ignores whether the overall level is too high or too low.
| variant | R MAE | R ρ | RBI MAE | RBI ρ |
|---|---|---|---|---|
| guess (shipped) | 9.659 | 0.8816 | 12.982 | 0.8150 |
| rate_only (slot term deleted) | 9.659 | 0.8816 | 11.832 | 0.8150 |
| **observed slot, as built** | 10.103 | 0.8751 | 12.548 | 0.8040 |
| observed slot plus strip step | 9.482 | 0.8872 | 12.002 | 0.8136 |
Using real slots moved rank correlation the wrong way for both categories: R −0.0064, RBI −0.0110. In the framing most generous to the layer, using slot data from the same season being projected, runs got worse in all four seasons. Strictly forward, using the prior season's slot, also 0 of 4. It lost the fight it was allowed to rig.
Look at the first two rows. Deleting the slot term entirely produces identical ρ for both categories and identical runs MAE to three decimals. That is an identity, not luck.
Every pair in this sample required 250 plate appearances the prior season, so the guessed slot never leaves the 3-to-5 bucket: slot 3 for 849 players, slot 4 for 64, slot 5 for 6. The runs multiplier in that bucket is 1.00. So for runs the slot adjustment is arithmetically a no-op. The RBI multiplier there is 1.075, applied to essentially everyone, which is a flat 7.5 percent inflation that reorders nobody. Hence the identical ρ.
Dropping the RBI term appears to gain 1.151 in MAE. Do not act on that. It moves zero ranks, which means it is purely a level shift, and the test rig here has no step that corrects the overall level while the real engine does. A uniform 7.5 percent RBI inflation is exactly what a calibration step absorbs. That 1.151 measures a hole in the test, not a flaw in the projections. Trust ρ. On ρ, real slots lose.
The layer's own RBI gain was 0.434. Deleting the constant gains 1.151. So the part you can attribute to slot accuracy is −0.716. It subtracts.
The reason is double-counting. Last season's rate stats already contain last season's batting order effect. Applying next season's multiplier on top counts it twice. With the degenerate guess that is harmless, because a constant cannot double-count anyone differently from anyone else. Give it real, varying slots and the double-count becomes real and uneven. We also tested and killed the tidier theory that runs broke because accurate slots activate a badly signed 6-to-9 constant: the re-estimated value made it worse, −1.194 versus −0.825, and neutralizing it to 1.00 did not help at −0.875.
One variant improved a level-proof metric: divide out last season's slot effect before applying next season's. It gained R ρ +0.00566 against both baselines, positive in 4 of 4 held-out seasons, with a confidence interval of +0.00326 to +0.01002, and a shuffled control that inverts the gain to −0.00485. RBI unchanged. That passed a pre-registered gate.
It is also tiny, roughly 0.6 percent, in-season only by construction, and unvalidated inside the live engine. This does not prove the RBI 1.075 should be removed. It does not prove the strip version should ship. It says nothing about plate appearances, pitchers, or true in-season projection. And 54 percent of pairs changed multiplier bucket, so the harm was not confined to a fringe.
Nothing changed in production. The observed-slot layer stays off, and both routes to reviving it are now closed.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.