We tested it, it did not survive the evidence, and the projections did not change. Here is the whole result, including what it does not prove.
There is a claim you have probably heard: reliever ERA is noise. The version that set this off put numbers on it, saying reliever ERA carries essentially no year-to-year signal (R² of 0.003, where R² is the share of next year's variation explained by this year's) while SIERA carries 0.13. Value relievers on process, not results. It is a reasonable thing to believe, and our projection system already leans that way. What it had never checked is whether one setting, tuned mostly on starters, was quietly under-correcting relievers. That was the question. The answer turned into something more interesting.
Building pairs of consecutive seasons from 2015 to 2025, requiring 120+ outs (40+ innings) in both years, gives 2,321 pairs: 1,206 starters, 745 relievers, 370 swing arms, with role assigned from the first year so nothing uses information a projection would not have.
| bucket | ERA → ERA R² | xERA → ERA R² | xERA → xERA R² |
|---|---|---|---|
| SP (n=1,206) | 0.0396 | 0.0853 | 0.1571 |
| RP (n=745) | 0.0054 | 0.0337 | 0.0618 |
| swing (n=370) | 0.0136 | 0.0470 | 0.0845 |
Reliever ERA is about 7× less self-predictive than starter ERA, and 0.0054 sits close to the outside claim's 0.003. That part replicates cleanly.
The headline 40× estimator advantage does not. That figure compares ERA predicting ERA against SIERA predicting SIERA, which is a same-metric comparison and says nothing about forecasting outcomes. The comparable ratio here is about 11×. And the number that actually matters for a projection, the estimator against next year's real ERA, is 0.0337 for relievers. Better than ERA. Still close to useless in absolute terms. The defensible sentence is: reliever ERA is not well predicted by anything in this data. Role churn is not the culprit either, since 87.2% of relievers who survive into the panel are still relievers the next year.
Our system nudges projected earned runs toward or away from observed ERA using a strength setting shipped at 10%. Testing that formula on its own, and comparing against the shipped 10% with a paired bootstrap (resampling the same pitchers thousands of times to get an error range), plus leave-one-season-out testing where the strength is picked on other seasons and scored on the held-out one:
| bucket | 0% vs 10% | LOSO pick | LOSO Δ MAE |
|---|---|---|---|
| POOLED | −0.0048 [−0.0110, +0.0017] ns | 0% in 8/10 | −0.0036 |
| SP | +0.0012 [−0.0085, +0.0108] ns | 5% in 10/10 | −0.0021 |
| RP | −0.0099 [−0.0192, −0.0012] SIG | 0% in 8/8 | −0.0103 |
| swing | −0.0141 [−0.0311, +0.0036] ns | 0% in 8/8 | −0.0169 |
Negative means the candidate beat shipped. The pooled verdict of "not significant" is a null starter effect and a significant reliever effect averaging into mush. Best strength: 10% for starters, 0% for relievers. The effect is small: for relievers it moves earned-run MAE from 6.709 to 6.652, about 0.06 earned runs over a season, with rank correlation improving from 0.1837 to 0.1976.
Before recommending a change, we checked whether anything else in the pipeline was already working on the same signal. Three separate corrections exist, not one. L1 multiplies projected earned runs toward actual ERA. L2 blends projected ERA 70/20 toward current-year xERA. L3 adjusts projected points toward xFIP, falling back to SIERA. Only L1 had ever been tested.
On a settled board of 683 pitchers, L1 fires on 661 (96.8%), L2 on 97 (14.2%), L3 on 430 (63.0%, xFIP every time, the SIERA fallback never fires). L1 and L3 fire together on 425 pitchers, 62.2%. All three on 81, 11.9%.
Then the split by whether a pitcher's ERA beat his xERA:
| Name | lucky (ERA < xERA) | unlucky (ERA > xERA) |
|---|---|---|
| L1 | +2.89 pts (n=360) | −3.23 pts (n=301) |
| L3 | −4.07 pts (n=234) | +4.37 pts (n=196) |
They are exact opposites. L1 rewards the pitcher who got lucky; L3 punishes him. On the 425 co-firing pitchers the signs oppose in 352 cases, 82.8%, and 62.6% of the combined magnitude cancels: mean |L1| of 3.25 plus mean |L3| of 4.22 gives 7.47 points of gross adjustment collapsing to a mean absolute net of 2.80. Mean net is −0.02 points, median −0.52.
What survives is close to a coin flip: 221 pitchers (52.0%) end up moved in the estimator direction, 181 (42.6%) in the opposite direction, 23 (5.4%) under a quarter point. The largest surviving distortions, where the cancellation happened to fail, are Garrett Crochet at net +10.3, Cal Quantrill +8.7 and Hunter Greene +8.2.
One thing here is objectively wrong and now fixed: the note attached to L1 described it as buying the unlucky pitcher and selling the lucky one. The code does the opposite. L3 is what implements that stated intent. Whether L1's behavior is wrong is genuinely arguable, since the base it multiplies is skill-derived and pulling a skill projection partway toward results is a defensible idea under a different name. The description was still backwards, and someone tuning it from that description would have moved it the wrong way.
We committed in advance to the thing that would settle it: turn the layers off in the full pipeline and see whether the board or the realized accuracy actually moves. If it does not, this is cosmetic, not a defect.
It does not. Accuracy, on 1,764 pairs with a fresh rebuild for each of the 8 on/off combinations:
| variant | MAE | Δ vs shipped | 95% CI | Pearson | Spearman |
|---|---|---|---|---|---|
| shipped | 77.747 | — | — | .4675 | .4133 |
| L1 off | 77.646 | −0.102 | [−0.289, +0.076] ns | .4704 | .4150 |
| L3 off | 77.927 | +0.180 | [−0.010, +0.382] ns | .4621 | .4078 |
| L1+L3 off | 77.823 | +0.076 | [−0.165, +0.321] ns | .4639 | .4107 |
Every interval includes zero. The direction is consistent across all three metrics: removing L1 helps slightly, removing L3 hurts slightly. Both at roughly 0.2% of MAE. The idea survives in sign and dies in size.
Ranks were re-run after we found and froze a source of non-reproducibility in the board loading path, using 2 loads per variant, all 5 variants byte-identical:
| variant | mean abs Δrank | max | ≥1 | ≥10 | ≥25 | Spearman vs shipped |
|---|---|---|---|---|---|---|
| L1 off | 0.76 | 9.4 | 147 | 0 | 0 | 0.999625 |
| L2 off | 0.58 | 22.7 | 78 | 9 | 0 | 0.999583 |
| L3 off | 0.30 | 11.2 | 47 | 1 | 0 | 0.999868 |
| L1+L3 off | 0.92 | 9.1 | 179 | 0 | 0 | 0.999522 |
No pitcher moves 25 ranks under any variant. Mean displacement is under one rank. The earlier claim of "Spearman ≥ 0.9996" does not reproduce; the correct bound is ≥ 0.9995, and we state the corrected figure rather than defend the old one. The withdrawn numbers were essentially right: the largest change in mean absolute rank movement across variants was 0.039 ranks.
L2 is the exception. Re-validated on the current pipeline, turning it off costs +0.1925 in ERA-MAE, with a range of [+0.1375, +0.2489]. L2 earns its place and should not be touched.
The reliever finding does not survive the full pipeline. Killing L1 helps starters more (−0.263 on a 101.25 MAE base) than relievers (−0.031 on 67.35), in both absolute and relative terms, the reverse of the isolated result. Neither is significant. The honest reading is that the reliever optimum was a product of testing one pathway alone, and it is withdrawn as actionable. The raw persistence asymmetry stands as a fact about pitchers; the tuning recommendation does not.
The reliever panel requires 40+ innings in both seasons. A reliever who loses his job and throws 30 innings drops out entirely, so the sample skews toward role-stable arms and, if anything, understates how unpredictable real relievers are. The co-firing measurements come from a snapshot of today's player pool, not a backtest; signs are exact, magnitudes approximate, because later normalization steps can rescale them.
We also had to withdraw a number. We initially reported that killing L1 moved final board points by 0.45 against a gross adjustment of 3.25, roughly 7× attenuation. That was measured by differencing cold loads, and the cold-load path turned out to be non-deterministic, with drift rather than random jitter across identical runs. The 7× is unmeasured, not merely imprecise. The lesson behind it, that a correction's size where it is applied is not its size on the board you actually read, is plausible and unestablished.
Still open: which direction the net ERA-luck treatment should point, whether the 52/43 residual split is signal or noise, whether L2 interacts with the other two, and whether the reliever result holds without the innings gate.
Nothing. No production change is warranted. L1 and L3 have no accuracy or rank case for removal; L2 is actively validated and its shipped shares are optimal. The description error is corrected. We did not ship a role-conditional branch for relievers: it buys roughly 0.004 MAE over the simpler global setting, which does not clear the bar for adding complexity.
The interim suggestion to drop L1's strength from 0.10 to 0.05 is retracted. It was measured in isolation, and live, most of it gets cancelled downstream. If anyone wants to revisit this, the order is: decide the intended direction first, test the surviving layer in the full pipeline, re-run the starter/reliever split against the composed system rather than one pathway, and only then touch a strength.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.