Six signals that sound like edges, tested on seasons the model had never seen, and failed. Publishing the nulls is the only way the wins mean anything.
Almost every public analytics piece reports something that worked. That is a selection effect, and it quietly makes every model look smarter than it is. So here are the signals we tested, believed in, and threw away. Each one was a real hypothesis with a real mechanism, and each one failed a test we wrote down before we ran it, on seasons the model had never seen.
Our rule: a ranking input has to earn its way in on held out seasons, not on a plausible story or a single before and after screenshot. Six candidates below did not earn in. They are not in the model.
This is the one we most expected to work. Velocity, movement, spin axis, release point: the whole Statcast arsenal, folded into a pitcher's fantasy projection. The theory is obvious. Better stuff should mean better results.
Tested with the pitcher's role held constant, using leave one fold out validation, it helped in 0 of 4 folds. The confidence interval was tight enough to exclude even a half percent gain in our average miss. Not "we could not detect it." Closer to "if an effect is there, it is smaller than half a percent."
The reason is that role already carries the information. Managers do not hand high leverage innings to pitchers with bad stuff. By the time you know a pitcher is an ace, a mid rotation starter or a middle reliever, the shape data has told you what it has to tell. We ran the same family five more ways, including within season and per pitch versions. All rejected.
Spin is real. It is one of the better understood physical inputs in baseball, and it genuinely predicts whiffs.
It still failed as a ranking input, and the distinction is the entire point. Fantasy scoring is an accumulation of innings, runs, wins and strikeouts. A spin driven strikeout edge that comes with fewer innings is close to a wash in points formats. Spin is rejected as a standalone ranking input, not as a baseball fact, and one narrow subgroup effect did survive. We kept the finding at the size the evidence actually supports.
We joined body measurements to dynasty value on a fully ID keyed match, so this was not a name matching failure or a coverage problem. It changed how much of dynasty value the model can explain by −0.00000.
Bloodline for hitters, sprint speed and raw exit velocity were rejected in the same study. That last one surprises people most, so it is worth being precise: exit velocity is not useless, it is already inside the projection through the median outcomes it produces. Adding it again as a separate term is double counting, and double counting shows up as noise.
We built five independent attempts at spotting a breakout before it appeared in results, using different data every time, including a Statcast arm and a pitcher specific arm.
The cleanest of them was negative in 8 of 8 seasons at both checkpoints. That is not a noisy null that a bigger sample would rescue. What our detectors actually do is identify breakouts as they happen, which feels like foresight and is not. There is a real gap between a model that recognizes a change quickly and one that sees it coming, and we could not cross it.
Dynasty formats reward young players, so a thin sample youth premium sounds like free value: a 23 year old with 400 major league plate appearances should carry more forward value than his production alone suggests.
Tested on seasons the model had never seen, the premium did not get earned more reliably by the players it was supposed to help. It is refuted and gated off. The related idea, that we should warm our dynasty ranks toward market consensus, failed the same way: every lever that moved us toward the market made the forward looking backtest worse, and the further we moved, the worse it got.
Blending scouting future value grades into prospect rankings failed on classes the model had never seen. The editorial top 100 rank alone out predicted the grades in every class we tested.
| Signal | How it was tested | Result |
|---|---|---|
| Pitch shape, levels | Role controlled, leave one fold out | Helped in 0 of 4 folds |
| Spin rate percentile | Standalone ranking input | Rejected, one subgroup survives |
| Height, weight, BMI | ID keyed join to dynasty value | Explaining power moved −0.00000 |
| Early breakout detection | Five arms, pre registered | Negative in 8 of 8 seasons |
| Thin sample youth premium | Held out dynasty seasons | Refuted, gated off |
| Scouting tool grades | Held out prospect classes | Failed, editorial rank wins |
Two reasons, and neither is modesty.
The first is that a null result is a real finding. Knowing that pitch shape adds nothing once role is known is worth as much as any positive result, because it tells you where not to spend your attention.
The second is calibration. A model that only reports its hits cannot be evaluated by anyone, including the people building it. We keep an audit record of every one of these, including the ones where our own first answer was wrong and we had to withdraw it. If we tell you a signal works, this page is the reason that claim should carry weight.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.