We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
A reasonable belief: if a prospect-ranking formula built out of hand-picked weights adds real signal, then letting the data pick those weights should add more. Our minor league composite scores batting prospects using three inputs, weighted 0.15 for OPS, 0.10 for level reached, and 0.05 for strikeout rate. Those numbers were chosen a priori, meaning before anyone looked at how they performed. That earlier test found the composite added +0.0402 to Spearman rank correlation, a measure of how well the ordering of prospects matches the ordering of their eventual careers, over a pedigree-only baseline. We declined to call it shippable partly because someone could always say the weights were guessed. So we fit them properly. The guesses won.
Nine cohort years. For each one, the three weights were fitted on the other eight years only, then the held-out year was scored using those weights and never touched its own fit, not its rows and not its scaling. That is leave-one-year-out testing, and it is the honest version: the model is graded on seasons it never saw. We searched for the weights that maximized rank correlation directly rather than minimizing squared error, because squared error on a target where most prospects amount to nothing would chase the handful of enormous careers and wreck the ordering, which is the thing we actually care about.
Average out-of-sample correlation came in at +0.319 for pedigree alone, +0.359 for the a-priori weights, and +0.373 for the fitted ones. Fitted beat pedigree by +0.0535, with a 95 percent confidence interval of [+0.0025, +0.1047] and 7 of 9 years positive. The a-priori weights beat pedigree by +0.0402, interval [+0.0048, +0.0752], 6 of 9 years. But fitted over a-priori? +0.0133, t of +0.58, interval [−0.0300, +0.0532], 6 of 9 years positive. That interval contains zero. The pre-declared bar required fitting to beat both baselines and to exclude zero against the guesses. It cleared the first and failed the second. Fitting is rejected.
Look at what the fitted weights actually did across folds. The OPS weight ranged from +0.06 to +0.39, a spread of 0.33 and a 6.5-fold swing: 0.06 when 2015 was held out, 0.39 when 2016 was. Level weight was stable at +0.07 to +0.11. Strikeout weight ran +0.20 to +0.35. All three were positive in all nine folds, so the direction of each input is real. The magnitudes are not. No one should quote a single fitted weight vector as the answer, because the folds do not agree on one.
Here is the interesting wrinkle for anyone who revisits this. The fit consistently wanted far more strikeout-rate weight than we guessed, 0.20 to 0.35 against 0.05, a 4 to 7 times difference, and more OPS weight, roughly 0.25 against 0.15. And despite those much larger weights, the out-of-sample gain was about 0.013. Many different weight combinations score nearly the same. That flatness is exactly why the fit wobbles and why optimizing buys nothing.
The earlier writeup listed as a limitation "a join covering only 22–56 batters per year." That phrasing implied thin coverage and was wrong in effect. Those are absolute counts, and the cohorts only contain 25 to 60 batters to begin with. Measured coverage is 413 of 460, or 89.8 percent, ranging from 84.6 to 94.8 percent by year. That has been corrected.
We also checked whether the 150-plate-appearance floor quietly drops a biased slice of players. Mean realized career outcome was 1012 for the 413 joined batters against 1044 for the 47 dropped, a ratio of 1.03 where 1.00 means the filter is outcome-neutral. With only 47 dropped players, that check can rule out a large effect and not a moderate one, and single years swing hard on 3 to 8 players: the 2017 dropped mean was 79, the 2011 dropped mean was 1960. The concern is substantially reduced, not formally retired.
It does not say the minor league layer should ship. It removes the "you guessed the weights" objection and tightens the estimate. It does not connect anything to the engine, and the governance rules around that are unchanged. It establishes no particular weight vector; if this layer is ever turned on, use the a-priori weights, not because they are optimal but because they are the only ones not chosen by looking at this data. It says nothing about pitchers, since the composite is batting-only. And it does not show the roughly +0.05 edge survives a different horizon. The earlier work found the minor league signal decays as you extend the evaluation window, +0.108 at three years falling to −0.017 at twelve. This run is seven-year only.
The falsifier is clear enough. Run the same leave-one-year-out protocol at a three-year and a twelve-year horizon. If fitted weights still fail to separate from the guesses at both ends, the flat surface is a property of the signal and not of this window, and the a-priori numbers are settled. If fitting suddenly matters at three years, the short horizon is where the real information lives and everything above is scoped too narrowly.
Evidence strength: moderate. Nine cohort years and a paired bootstrap, meaning we resampled the same years repeatedly to see how much the gap moves. Enough to reject the fitted weights. Not enough to ship anything.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.