We Tried to Optimize Our Minor League Prospect Model. Guessing Beat It.

We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.

A reasonable belief: if a prospect-ranking formula built out of hand-picked weights adds real signal, then letting the data pick those weights should add more. Our minor league composite scores batting prospects using three inputs, weighted 0.15 for OPS, 0.10 for level reached, and 0.05 for strikeout rate. Those numbers were chosen a priori, meaning before anyone looked at how they performed. That earlier test found the composite added +0.0402 to Spearman rank correlation, a measure of how well the ordering of prospects matches the ordering of their eventual careers, over a pedigree-only baseline. We declined to call it shippable partly because someone could always say the weights were guessed. So we fit them properly. The guesses won.

How the test was run

Nine cohort years. For each one, the three weights were fitted on the other eight years only, then the held-out year was scored using those weights and never touched its own fit, not its rows and not its scaling. That is leave-one-year-out testing, and it is the honest version: the model is graded on seasons it never saw. We searched for the weights that maximized rank correlation directly rather than minimizing squared error, because squared error on a target where most prospects amount to nothing would chase the handful of enormous careers and wreck the ordering, which is the thing we actually care about.

Average out-of-sample correlation came in at +0.319 for pedigree alone, +0.359 for the a-priori weights, and +0.373 for the fitted ones. Fitted beat pedigree by +0.0535, with a 95 percent confidence interval of [+0.0025, +0.1047] and 7 of 9 years positive. The a-priori weights beat pedigree by +0.0402, interval [+0.0048, +0.0752], 6 of 9 years. But fitted over a-priori? +0.0133, t of +0.58, interval [−0.0300, +0.0532], 6 of 9 years positive. That interval contains zero. The pre-declared bar required fitting to beat both baselines and to exclude zero against the guesses. It cleared the first and failed the second. Fitting is rejected.

The instability is the real finding

Look at what the fitted weights actually did across folds. The OPS weight ranged from +0.06 to +0.39, a spread of 0.33 and a 6.5-fold swing: 0.06 when 2015 was held out, 0.39 when 2016 was. Level weight was stable at +0.07 to +0.11. Strikeout weight ran +0.20 to +0.35. All three were positive in all nine folds, so the direction of each input is real. The magnitudes are not. No one should quote a single fitted weight vector as the answer, because the folds do not agree on one.

Here is the interesting wrinkle for anyone who revisits this. The fit consistently wanted far more strikeout-rate weight than we guessed, 0.20 to 0.35 against 0.05, a 4 to 7 times difference, and more OPS weight, roughly 0.25 against 0.15. And despite those much larger weights, the out-of-sample gain was about 0.013. Many different weight combinations score nearly the same. That flatness is exactly why the fit wobbles and why optimizing buys nothing.

A correction on coverage

The earlier writeup listed as a limitation "a join covering only 22–56 batters per year." That phrasing implied thin coverage and was wrong in effect. Those are absolute counts, and the cohorts only contain 25 to 60 batters to begin with. Measured coverage is 413 of 460, or 89.8 percent, ranging from 84.6 to 94.8 percent by year. That has been corrected.

We also checked whether the 150-plate-appearance floor quietly drops a biased slice of players. Mean realized career outcome was 1012 for the 413 joined batters against 1044 for the 47 dropped, a ratio of 1.03 where 1.00 means the filter is outcome-neutral. With only 47 dropped players, that check can rule out a large effect and not a moderate one, and single years swing hard on 3 to 8 players: the 2017 dropped mean was 79, the 2011 dropped mean was 1960. The concern is substantially reduced, not formally retired.

What this does not prove, and what would change it

It does not say the minor league layer should ship. It removes the "you guessed the weights" objection and tightens the estimate. It does not connect anything to the engine, and the governance rules around that are unchanged. It establishes no particular weight vector; if this layer is ever turned on, use the a-priori weights, not because they are optimal but because they are the only ones not chosen by looking at this data. It says nothing about pitchers, since the composite is batting-only. And it does not show the roughly +0.05 edge survives a different horizon. The earlier work found the minor league signal decays as you extend the evaluation window, +0.108 at three years falling to −0.017 at twelve. This run is seven-year only.

The falsifier is clear enough. Run the same leave-one-year-out protocol at a three-year and a twelve-year horizon. If fitted weights still fail to separate from the guesses at both ends, the flat surface is a property of the signal and not of this window, and the a-priori numbers are settled. If fitting suddenly matters at three years, the short horizon is where the real information lives and everything above is scoped too narrowly.

Evidence strength: moderate. Nine cohort years and a paired bootstrap, meaning we resampled the same years repeatedly to see how much the gap moves. Enough to reject the fitted weights. Not enough to ship anything.

All articles · Redraft rankings · Dynasty rankings

Skip to main content

Redraft Rest of Season

Sorted by -
Loading history…
Official MiLB Prospect Rankings

Official MiLB Prospect Rankings

Loading…
Loading rankings…
Overview

Operations Dashboard

Portfolio and league analytics.
Workspace ready
Use Sync to import a roster.

2026 FYPD Class

Players entering the MiLB system for the first time: the 2026 MLB Draft class and first-time international signees. Ranked by dynasty value: a measured expectation from draft slot or debut production, adjusted by the age curve. For established prospects, see the Prospect board.

Prospect Player Rankings

Top-100 + all 30 org Top-30 lists · ranked by projected Prospect Value · AAA Statcast tools →
Loading prospect rankings…
Loading study…

Methodology

Every signal, formula, and data source the model uses for player evaluation - organized by product.

Loading methodology…

Risers & Fallers - last 30 days