We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
Here is a thing that sounds obviously true: if you want to know how much of the rest of the season an injured player will actually be available for, you should look at what he hurt. An elbow is not a hamstring. A 60-day designation is not a 7-day one. Our system does not use any of that. It applies a flat 0.6 availability factor to every player on the injured list, regardless of body part, diagnosis or designation. That is a placeholder, and it deserved to be challenged.
So we built three replacements and scored them against it. All three lost.
Each estimator predicts rest-of-season points for players on the IL, and we compare that to what they actually produced. The scoring number is median absolute error, meaning the typical miss in points, with lower being better. Everything was fitted only on 2019, 2021 and 2022, then tested on 2023 through 2025, seasons the model had never seen. We also scored three separate as-of dates, because a fix that only works in one week of the calendar is not a fix.
The bar was set in advance: beat the flat 0.6 by at least 10 percent, and do it consistently.
The conditional model, the most sophisticated arm, beat production by 9.2 percent at the July 31 primary date and 10.5 percent at August 31. At June 15 it was 23.5 percent worse than doing nothing. The simple per-designation median looked better at July 31 at plus 13.9 percent and a startling plus 85.7 percent at August 31, then went to minus 11.1 percent in mid-June. Using the designation number alone was the worst of the lot, between 14.9 and 36.4 percent worse than production at every date. Sample sizes ran from 118 held-out players in June to 350 in late August.
Verdict: rejected. Nothing changes in production.
The usual explanation for a model that works in training and fails in testing is overfitting, where the model memorizes noise from the data it learned on. That is not what happened here. We scored the conditional model on its own fitting seasons and it came in at minus 2.9 percent, worse than a flat constant on the very data it was built from. The features do not carry the signal. Full stop.
What makes this genuinely uncomfortable is that the fitted table is medically excellent. It learned, from data alone, that a 60-day designation with Tommy John means 226 days, that a 60-day elbow means 155 days, that a 7-day concussion means 11 days and a 10-day hand means 11 days. A twentyfold spread between Tommy John and a concussion, nobody hand-coded it, and every line of it is correct.
It still does not help. Knowing the typical timeline for a diagnosis does not tell you where one specific player sits inside a range whose middle half spans months.
The 0.6 is not well calibrated. It contains no injury information whatsoever. It wins because it sits near the middle and is boringly wrong, while the alternatives are confidently wrong at both ends. Designation number is far too short. Conditional medians are too long for everyone who comes back early. A constant near the average beats a confident wrong answer. That is a statement about the competitors' volatility, not about 0.6 being right.
Worth knowing how much is left on the table: an oracle with perfect knowledge of the actual return date beats production by 32.6, 77.9 and 100 percent across the three dates. The ceiling is real and it is high. Three estimators failed to reach any of it.
It does not prove no estimator can work. It proves these three, over these fields, do not. It does not prove 0.6 is correct. It is only the best of what was tested, and roughly 78 percent of mid-season availability error is still unclaimed.
The answer is not a cleverer model over the same information. It is a data source carrying an actual expected return date, and we went looking. None exists that we can reach. The live return date is just derived from the designation. One retrospective date field simply mirrors the transaction date. Of 14 live comments containing a Month DD token, 13 refer to events already past.
One forward-looking signal does arrive: rehab-stage language, present on 97.5 percent of live comments. We cannot test it, because our historical transaction records carry no comment field. So we start archiving daily injury snapshots now, and the question becomes answerable in about a season.
Until then, the 0.6 stays. The difference is that it is now a measured choice with a known ceiling above it, instead of a default nobody had ever scored.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.