Does pitch SHAPE predict the model's error?

We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.

_This is one of a series on hypotheses that failed. Every projection change here has to beat a held-out test before it ships, and most candidates do not. Publishing only the ones that worked would misrepresent how the model is built._

What we thought might be true

the analysis landed this morning and ingested the analysis through 2026: 10,788 pitch rows over 3,652 pitcher-seasons, 9 pitch types, 100% coverage on velocity, induced vertical break, horizontal break, and spin. Its commit message states the purpose plainly, that pitch design had been deferred by data rather than by evidence, because pitch_arsenal_<year>.json carried 0.0 for velocity and spin in every vintage.

So: with real movement data, does pitch shape predict where the projection model is wrong?

Audit classification (initial hypothesis): Valuation

How we tested it

the analysis (committed in a pinned build) ran the walk-forward and reported "pitch shape ADDS NOTHING beyond a constant bias correction" (held-out average error 67.8 with shape vs 66.3 bias-only). Re-running it reproduces that exactly.

The conclusion is correct. The test, as run, could not have detected the effect it was looking for:

1. The target is volume, not quality. The residual was actual - proj in total fantasy points, and 47% of that residual's variance is innings error (corr(resid, actualIP - projIP) = 0.685, corr(resid, actualIP) = 0.644). Pitch shape cannot predict how many innings a pitcher throws; that is health and role. The MVP Machine's claim is about the quality of stuff, which lives in rate. Testing shape against a volume-dominated total buries the signal by construction. 2. A role proxy was the strongest "shape" feature. n_shapes, the count of pitch types clearing 50 pitches, had the most sign-stable correlation of any feature (+0.220 / +0.119 / +0.140 / +0.103). It is not a shape. It correlates with realized innings at +0.53, +0.49, +0.46, +0.42 across the four folds. It is a starter-versus-reliever flag wearing a shape costume. 3. One held-out fold. Fit on 2022 to 2024, tested on 2025 alone (n=286). A single fold cannot separate "no signal" from "unlucky season." 4. No uncertainty. 67.8 versus 66.3 was reported without a confidence interval, so it was unknown whether shape lost by noise or by a real margin.

Stage 2 fixes all four: a rate target, feature sets that isolate the role proxy, leave-one-fold-out over all four outcome seasons, and a 2,000-replication bootstrap on the average error difference.

What actually happened

Bar throughout is deliberately hostile: beat the baseline on held-out average error. Correlating with the residual is not enough.

Target 1: total points (stage 1's target), baseline = constant bias

Feature setaverage error biasaverage error +featuresΔ95% confidence intervalfolds helped
all (stage 1)76.55575.719−0.836[−1.863, +0.252]3/4
shape only76.55576.363−0.192[−0.963, +0.528]3/4
movement only76.55576.280−0.275[−0.914, +0.322]4/4
velocity only76.55576.528−0.026[−0.471, +0.412]2/4
n_shapes only76.55575.798−0.757[−1.512, +0.025]3/4

Every confidence interval straddles zero. Stage 1's null now has power behind it.

Target 2: rate (points per IP), baseline = constant bias

Feature setaverage error biasaverage error +featuresΔ95% confidence intervalfolds helped
all (stage 1)0.8500.782−0.068[−0.088, −0.047]4/4HELPS
shape only0.8500.825−0.025[−0.038, −0.014]4/4HELPS
movement only0.8500.831−0.018[−0.028, −0.008]4/4HELPS
velocity only0.8500.848−0.002[−0.008, +0.004]0/4null
n_shapes only0.8500.785−0.064[−0.083, −0.045]4/4HELPS

Switching to a rate target appears to resurrect the signal: movement helps in all four folds with a confidence interval excluding zero. Note that n_shapes alone recovers most of the full model's gain, which is the tell.

Target 3: rate, baseline = bias **plus role** (`projIP`, SP flag, `engineRoleRet`)

Feature setaverage error bias+roleaverage error +featuresΔ95% confidence intervalfolds helped
all (stage 1)0.6660.669+0.003[−0.001, +0.007]0/4null
shape only0.6660.670+0.003[−0.001, +0.007]0/4null
movement only0.6660.668+0.002[−0.000, +0.003]0/4null
velocity only0.6660.666+0.000[−0.001, +0.001]0/4null
n_shapes only0.6660.666−0.000[−0.001, +0.000]0/4null

The stage-2 signal was role confounding, entirely. Once the baseline knows whether a pitcher is a starter and how many innings the engine projects for him, pitch shape adds nothing: 0 of 4 folds for every feature set, and n_shapes collapses to exactly zero, which is what a pure proxy does when its target is controlled. The movement-only confidence interval, [−0.000, +0.003] on a base of 0.666, rules out even a 0.5% average error improvement. This is a powered null, not an absence of evidence.

projIP is the engine's own forward innings projection, so using it as a control is leak-free: it is known at projection time.

Target 4: shape CHANGE, which is what "pitch redesign" actually means

Levels answer "is his stuff good." The book's claim is about changing stuff. Features are year-over-year deltas from T−1 to T, baseline bias+role, outcome seasons 2023 to 2025:

Feature setaverage error bias+roleaverage error +featuresΔ95% confidence intervalfolds helped
delta only0.6340.630−0.004[−0.012, +0.004]2/3
delta + level0.6340.636+0.003[−0.007, +0.013]1/3
**velocity gain**0.6340.634+0.000[−0.002, +0.002]1/3
**ride gain** (iVB)0.6340.634−0.000[−0.003, +0.003]2/3

Also null. Velocity gain and ride gain, the two canonical Driveline interventions, land at exactly zero. But this variant has 3 folds and n=617, and the delta only confidence interval reaches −0.012, about 1.9% of the base average error. It cannot exclude a small real effect, so it closes Deferred, no support found rather than Rejected.

Verdict

Rejected. Nothing in the projections changed as a result of this study.

What this does not prove

- That pitch design does not work in reality. This tests whether shape predicts managr's forecast error for next-season fantasy value. A redesign that works shows up in the following season's strikeouts and ERA, which the engine already reads through K%, xERA, and the comp blend. The finding is about incremental forecast value, not about baseball. - That shape change is useless. The delta test is underpowered (3 folds, n=617). - Anything about in-season use. A mid-season shape change detected in real time is a different question on a different horizon, and was not tested. - That the engine's role handling is correct. See below.

What would change our mind

The conclusion "do not add pitch shape to the projection" would be falsified by:

1. A shape-change test with power — extending the delta panel as 2026 and 2027 complete (the shape vintages start in 2021, so each new season adds a fold), driving the delta only confidence interval inside ±0.005 while the point estimate stays negative. 2. A per-pitch-level test rather than a pitcher-season aggregate. This study collapses an arsenal to primary-fastball shape plus two separation scalars. A model over the full per-pitch matrix might find structure the aggregate destroys. The data supports it; this harness does not attempt it. 3. Shape predicting a component the aggregate hides, most plausibly K% specifically, rather than composite points per IP.

_Written up from the internal audit record. The numbers are the ones the study produced; nothing was re-run to make the story cleaner._

All articles · Redraft rankings · Dynasty rankings

Skip to main content

Redraft Rest of Season

Sorted by -
Loading history…
Official MiLB Prospect Rankings

Official MiLB Prospect Rankings

Loading…
Loading rankings…
Overview

Operations Dashboard

Portfolio and league analytics.
Workspace ready
Use Sync to import a roster.

2026 FYPD Class

Players entering the MiLB system for the first time: the 2026 MLB Draft class and first-time international signees. Ranked by dynasty value: a measured expectation from draft slot or debut production, adjusted by the age curve. For established prospects, see the Prospect board.

Prospect Player Rankings

Top-100 + all 30 org Top-30 lists · ranked by projected Prospect Value · AAA Statcast tools →
Loading prospect rankings…
Loading study…

Methodology

Every signal, formula, and data source the model uses for player evaluation - organized by product.

Loading methodology…

Risers & Fallers - last 30 days