We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
The idea is old enough to be a book. Big Data Baseball built its case on the 2013 Pirates, and the insight was never "ground-ball pitchers are good" or "positioning is good." It was that the two interact: a sinkerballer in front of a well-positioned infield is worth more than the sum of the parts. Value the pitcher and the defense as one system.
Our projections currently treat defense as a hitter-only input. Statcast Outs Above Average feeds a small playing-time bump for good glove men, on the sensible theory that plus defenders rarely get benched, and it feeds a historical replay layer. No pitcher reads any of it. If the Pirates were right, that is a hole. So we tested two things in order: does the defense behind a pitcher predict his next-season run prevention beyond what his own skill signals already say, and is that effect bigger for ground-ball guys?
We took pitchers with at least 40 innings in back-to-back seasons, paired year H to year H+1, threw out 2020 on both sides for the obvious reason, and ended up with 1,321 pairs across holdout seasons 2018, 2021, 2022, 2023 and 2024. Team defense is the sum of fielding runs prevented across qualified fielders, z-scored within each season so league-wide scale drift and the 2023 rule change do not leak into the variable itself. Coverage is qualified fielders only, roughly 250 a season or 8.3 per team, so the measure is noisy. Noise pushes effects toward zero, which means everything below is if anything an understatement.
The baseline the defense term has to beat is a model using each pitcher's own year-H ERA, xERA, K%, BB%, barrel rate, hard-hit rate, age and innings. Out of sample, meaning on seasons the model never saw, that baseline explains 10.6% of next-year ERA, 49.2% of K% and essentially none of BABIP, 0.5%. Defense is measured against what is left over.
Per one standard deviation of team defense, next-season ERA moves −0.114 on its own. Add park factors and it is −0.089 (p=0.009, meaning a result this big would come up by chance about nine times in a thousand). Add our shipped catcher-framing layer and it is −0.077. Weight by innings and it is −0.103. Infield-only defense gives −0.077, slightly weaker than all fielders, which is awkward for anyone selling a shift-specific story, since shifts are an infield thing.
The design checks out. Park enters at +0.188 with the right sign and a believable size. Framing enters at +0.027, near zero and wrong-signed, so the framing layer we already ship is not secretly what this is picking up. And the negative control is clean: defense does not predict strikeout rate at all (r=0.007, p=0.79). It shouldn't. Fielders cannot catch a strikeout.
Here is the problem. Team defense is one number per team per season. Every pitcher on the 2024 Guardians shares it. So 1,321 observations are really 148 independent team-seasons, and the naive p-value of 0.0004 from the first pass was meaningless. Once you correct for that and add team fixed effects, which force the estimate to come only from a team's own year-to-year swings rather than from good organizations being good at everything, the coefficient halves to −0.041 with a range of [−0.142, +0.052] and p=0.32.
Two facts pull against each other. Sixty-six percent of the variation in team defense is within team year over year (the autocorrelation is only 0.354), so the fixed-effects test retains most of the signal and its null is not purely a power problem. But the smallest effect that test could reliably detect is 0.116, and the thing we are asking it to find is 0.089. The test is underpowered to detect the very effect it is testing, and its range contains the pooled estimate.
That means it rules out nothing in either direction. We cannot separate "this defense" from "this organization." That is the central open question, and pretending otherwise would be dishonest.
The ground-ball interaction, tested on the 770 pairs where ground-ball rate is available across 88 team-seasons:
| Term | β | 95% CI | p |
|---|---|---|---|
| defense_z | −0.053 | [−0.126, +0.020] | 0.155 |
| gb_z | −0.091 | [−0.169, −0.012] | 0.023 |
| **defense × GB** | +0.008 | [−0.076, +0.092] | 0.852 |
| park_z | +0.175 | [+0.105, +0.245] | <0.001 |
Zero, and to the extent it has a sign it is the wrong one: positive means the defensive benefit shrinks for ground-ball pitchers. Terciles do not line up either: low-GB −0.088, mid-GB −0.035, high-GB −0.081. That is noise, not a dose-response curve.
But we found no support for the interaction. We did not refute it. The smallest interaction this test could detect is 0.120 while the main effect is 0.053, so the range still allows an interaction as large as the main effect itself.
Ground-ball rate on its own looked like a nice free feature at −0.090 (p=0.024). It isn't. Add SIERA to the baseline and it collapses to −0.023 (p=0.564), because SIERA already contains ground-ball rate. Classic trap: a signal that looks new only because the baseline was too weak to already have it.
Running the same strong-baseline check on the defense coefficient dropped it from −0.089 to −0.056. We assumed SIERA. It wasn't. On the identical 770 rows, the baseline without SIERA gives −0.055 and with SIERA gives −0.056. The drop was the sample period.
| Outcome seasons | Regime | β | 95% CI | p | n |
|---|---|---|---|---|---|
| ≤ 2022 | shifts legal | −0.150 | [−0.319, +0.019] | 0.083 | 536 |
| ≥ 2023 | shifts restricted | −0.055 | [−0.131, +0.021] | 0.158 | 770 |
About 2.7 times larger before MLB restricted shifts. Which is exactly what the book and the rule change both predict: when you can put fielders anywhere, the defense behind a pitcher matters more. An edge that decayed, except this one was legislated away rather than competed away.
Suggestive, not established. The formal era difference is +0.091 with p=0.22 against a detection threshold of 0.208, so the two eras are not separable from noise here, and neither era estimate is individually significant. What we can say: the pooled −0.089 headline is partly borrowed from a rule regime that no longer exists, and the current-era number cannot be told apart from zero.
Other targets moved the way the mechanism says they should: BABIP −0.0026 (p=0.043) and WHIP −0.0137 (p=0.035).
Applied to the live 2026 pitcher board, 572 of 672 covered, at default scoring:
| Metric | This feature | park factors | hotColdScale |
|---|---|---|---|
| players affected | 441 | 1,057 | 418 |
| mean abs Δ | 2.5 pts | 6.0 | 9.0 |
| moved ≥ 10 ranks | 64 | — | — |
| moved ≥ 25 ranks | 2 | — | — |
| moved ≥ 50 ranks | 0 | 81 | 35 |
A top-60 starter swings 4.1 points on average and moves a median of one rank. The single biggest mover on the whole board is 13.8 points and 33 ranks. At the current-era coefficient of −0.055, shave another 38% off all of it. Both subsystems we have already examined are far bigger, and this one displaces nobody by 50 spots.
Research only. We are not shipping a pitcher-side defense term. Three things have to hold and only the first does: the association is real (yes), it is identified as causal rather than organizational (unknown, underpowered), and it is big enough to matter in the era we are actually projecting (no).
Two things would flip this. A fixed-effects estimate with real power, meaning enough seasons to push the detection threshold below about 0.06 while the coefficient stays near −0.09, would establish it as causal. Or a full engine-on/engine-off test showing the term improves walk-forward accuracy against the shipped projections rather than against this study's 10.6% baseline. A signal can survive a weak proxy and still be fully absorbed by the real model.
Going the other way, 2026 and 2027 landing near −0.05 with tighter ranges would confirm the call. That is the plan: rerun it after 2027, when the post-restriction era has four complete outcome seasons instead of two.
We tested nothing about relievers, nothing about in-season projection, and nothing about the hitter-side defense layer. And framing's near-zero coefficient here says nothing about whether our framing layer is calibrated correctly. It was a by-product of a test built for something else, on a partial sample.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.