A reader compared our dynasty starters to another public board and concluded we back proven production harder than the market. Under the full join, the conclusion flipped.
A reader sent over a careful comparison of our dynasty starting pitcher board against another public board. It listed thirteen starters we rank far above them, drew a reasonable conclusion from the pattern, and asked us to pressure test it. The conclusion was that our model believes more than the market does in established major league production becoming durable skill.
Every one of those thirteen observations was correct. The conclusion was still backwards, and the reason is worth more than the answer.
A "biggest disagreements" list is a tail by construction. If you read a philosophy off one, you will get the right answer only when the tail and the middle happen to agree. For a model that differs from the market in degree rather than in direction, they usually do not.
Instead of comparing hand picked names, we joined our whole dynasty board against a published consensus ranking and looked at the typical gap rather than the extremes. As of the August 2 board, 259 players matched.
| Group | Players | Median gap | We rank higher | We rank lower |
|---|---|---|---|---|
| Starting pitchers | 91 | 40 places colder | 31 | 59 |
| Everyone else | 168 | 39.5 places colder | 61 | 107 |
| All matched | 259 | 40 places colder | 92 | 166 |
We are colder than consensus on starting pitchers. Not by a little: on the median starter we sit forty places below where the market has him, and we rank almost twice as many starters lower than higher.
The thirteen names in the original comparison are real and they reproduce against this feed too. They are simply the far end of a spread whose middle points the other way. The same board that has us aggressive on a handful of arms also has a former Cy Young winner at our 352nd dynasty asset while consensus keeps him at 133.
One detail worth stealing for your own work. We ran this two ways: once against the top 500 players our board renders, and once against the full board.
The rendered top 500 gives a milder answer. That is not noise, it is structural. Cutting at 500 removes the deepest disagreements, and those are disproportionately players we rank far lower than the market, because a player we have at 480 and the market has at 120 is exactly the kind of row a top 500 cut can drop. Truncating the data moved the result toward the comparison's original conclusion.
Any time you compare two ranked lists, check whether your cut is symmetric. Usually it is not.
If the board is not rewarding proven workload, what is it doing? We checked, across every starter eligible arm on the board.
| Input | Correlation with dynasty value |
|---|---|
| Healthy rate of production | .85 |
| Projected innings | .61 |
| Projected fantasy points | .63 |
| Age | −.17 |
The dominant term is the rate a pitcher produces at when healthy. Innings matter, and matter clearly, but they are second.
The board hands us a natural experiment on that point. Three starters currently project for effectively the same workload, within one inning of each other, and we rank them nowhere near each other.
| Starter | Projected innings | Healthy rate | Our rank |
|---|---|---|---|
| Dylan Cease | 168 | 521 | SP2 |
| Logan Webb | 168 | 334 | SP26 |
| George Kirby | 169 | 241 | SP50 |
Same volume. Forty eight places between the top and the bottom. The entire spread is rate.
That is also the answer to a disagreement the original comparison flagged. Kirby is a command first starter the market likes considerably more than we do, and the write up explaining their side argued that command ages well. We are not arguing with that. We are pricing what he produces when he pitches, and on this board that number is less than half of Cease's at the same innings.
You can see the same thing at the very top from the other direction: our number one dynasty starter projects for 140 innings, while the arm we have seventh projects for 205. If this were a workload machine that ordering would be impossible. It is a quality machine that respects volume, which is a different thing.
The comparison also concluded that we punish age harder than the market. That is true at the old end and misleading as a general statement.
| Age | 21 | 24 | 26 | 28 | 30 | 32 | 34 | 36 | 38 |
|---|---|---|---|---|---|---|---|---|---|
| Multiplier | 1.14 | 1.08 | 1.04 | 1.00 | 0.95 | 0.90 | 0.79 | 0.68 | 0.57 |
From 28 to 32 the penalty is about two and a half points a year. After 32 it roughly doubles. Across every starter in the sample, age correlates with dynasty value at only −.17, because most starters sit between 24 and 31 where that band is narrow. One 38 year old carrying a 0.57 multiplier is not evidence of a general anti veteran tilt. It is evidence of a curve that gets steep late.
Being colder than the market on pitchers is a choice we have already tested rather than a position we defend on principle. We built the levers that would warm our dynasty ranks toward consensus and measured them against realized forward production. Every one of them made the forward looking backtest worse, and the further we warmed, the worse it got. A related idea, a premium for young players on thin major league samples, was refuted the same way and is switched off.
So the coldness stays. Not because we like it, but because every attempt to remove it cost accuracy on the only test that matters.
This measurement moves. Three days before this piece, the same join put the starter median at 62 places rather than 40, and showed starters running colder than everyone else. Today starters and non starters are almost identical.
The direction has been stable across every look: we are colder, and we rank far more starters below consensus than above it. The size of the gap is not fixed and should not be quoted as though it were. If you ever see that median hit zero or the split reverse, the claim in this article is dead, and that is the number to watch.
That is the same discipline the article is about. One snapshot supports a direction. It does not support a decimal place.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.