The two most repeated pieces of fantasy advice, tested against 1,704 hitter seasons. Neither one held up, and the popular version was the worst rule we measured.
Every August, two pieces of advice get repeated until they sound like facts. Sell the hitter whose numbers have run ahead of his contact quality, because it is about to come down. Buy the star having a down year, because he will be his old self.
We tested both against our own backtest. Neither one holds. One of them is backwards.
The hitters everyone says to sell beat our projections by 12 to 16 percent. And "he will be his old self" was the worst forecasting rule we measured.
The argument makes sense on its face. Take a hitter with a .330 wOBA and a .284 xwOBA. wOBA is what he actually produced. xwOBA is what his batted balls were worth, based on how hard and at what angle he hit them. When the first runs well ahead of the second, the story writes itself: he has been lucky, so sell before it corrects.
So we checked. We took 1,704 hitter seasons from our backtest, skipped 2020, and compared what our engine projected for each player against what he actually did the following year.
The biggest over-performers, the ones beating their contact quality by .030 or more, came in 16 percent under-projected. Not over. Under. They were the most under-projected group in the entire study. Every other group, from strong to mild to neutral to the hitters running behind their contact quality, landed within 4 percent.
The players everyone tells you to sell beat their forecasts by more than anyone else.
That is less strange than it sounds, once you stop calling it luck. Beating your expected stats is partly a skill you keep: bat control, speed, where you tend to hit the ball, the park you play in half the time. The "it will come down" story treats all of that as noise. A good chunk of it is not.
Then we tried to break our own result. If these hitters really do beat their projections, we should be able to get more accurate by pulling them down less. So we tested the whole range, from the full correction we normally apply down to none at all.
Nothing moved. Our average miss stayed at 84 points at every setting. The order we ranked players in did not shift either, holding at 0.589 on a scale where 1.0 would be perfect. Removing the correction entirely narrowed the gap from 16 percent under to 11 percent under and bought no accuracy at all.
That is the more interesting finding, and it is the one that settled the question. The lever is not just pointed the wrong way. It is not connected to anything. And pushing it the way instinct says to push it does real damage: adding more downward correction made the average miss worse at every step, from 76 up to 87.
The reason no single adjustment works is that "over-performer" is not one kind of player. It is the hitters who keep it up and the hitters who fall apart, sitting in the same bucket. No single multiplier can separate them, because on the day you have to decide, they look identical.
A separate test, run off a different trigger by a different setup, found the same thing from the other end. That group came in 12.2 percent under-projected. The players who best fit the classic lucky profile, below-average power and ordinary speed, came in 14.3 percent under.
One more detail makes this stronger rather than weaker. To show up in a test like this, a player has to still be playing the next season. The ones who wash out completely never enter the sample. So the players we can measure look better than the full group really is, which tilts the entire study toward correcting more, not less. Our engine still lands 12 percent under on a sample already tilted that way. There is no room to push the correction harder. Doing it would only widen a gap that already points the wrong direction.
The other half of the advice is about the established star having a down year. The market is wrong here too, in a more interesting way than simply being backwards.
We took hitters who had genuinely been good, a .350 wOBA or better, age 27 or under, and who then had a down year. Then we looked at what they did next, across every season from 2015 on.
They do rebound, but only part of the way. The typical player won back about 40 percent of what he lost. In points per 650 plate appearances: 428 before the down year, 343 during it, 388 after. That is a real recovery and worth owning. It is nothing like a return to form.
Then we tested three ways of forecasting those players against what actually happened. Lower is better, and the numbers are how far each rule missed by:
| What you assume | How far off it was | Verdict |
|---|---|---|
| "He will be his old self" | 0.0375 | Worst of the three |
| "The down year is who he is now" | 0.0270 | Better |
| Split the difference, 60/40 | 0.0236 | Best |
The popular answer was the worst one. Betting that a down-year star returns to his old level costs you more than assuming the down year is his new normal. And both lose to simply splitting the difference between them.
Which raises the obvious question. If the blend wins, why have we not built it?
We did build it, behind a flag, and measured it properly. Across three strengths, our average miss on the target group stayed flat at 72. On the players the change actually fired on, accuracy got worse. Every setting made the model worse overall, and the harder we pushed, the worse it got.
The diagnostic explains why. The typical player in that group is already projected almost exactly right: 1.024 on rate, where 1.000 would be perfect, and 1.003 on playing time. The 11.5 percent gap that looked so promising is an average being dragged around by a handful of extreme seasons, not a mistake the model is making across the group. Two things drive it, and a rate adjustment cannot reach either one: the occasional career year no model sees coming, and the injury and availability discount we apply further down the line.
There is also no way for the lever to tell the two stories apart. A down year is followed by a bounce or by a decline, and on the day you make the call, the input looks the same. One player in our sample got lifted from 421 to 460 by the change. He then produced 85.
We publish the finding and leave the engine alone.
That is less satisfying than shipping a lever, and it is the honest read of what we found. The correction already in the model is strong enough to remove the systematic over-projection. That is a specific claim, and it is not the same as saying the setting is perfect. But nothing we found suggests it is too weak, and three separate attempts to improve it all came back flat.
The short version for a reader is this.
Do not sell a hitter just because his expected stats trail his results. On our numbers, that player beats his projection more often than not.
Do not pay full price for a down-year star on the assumption he returns to his old level. He gets about 40 percent of the way back, and "he will be his old self" was the worst-performing rule we tested.
And treat anyone who tells you either one with confidence as someone who has not checked.
The down-year group is small: 45 players at the strict age cut, 79 if we loosen it. A sample that size can flip.
These tests measure accuracy and ranking. A different goal, auction dollars in a specific format, or a category league where the shape of your production matters more than the total, could land somewhere else. We have not tested that.
One thing is genuinely untested. During the season, our full-year number pulls the forward projection toward the pace a player is actually on, which partly undoes the correction described here. We cannot test it yet, because it needs game-by-game historical splits we do not carry. We think the direction is right, and we have not shipped a change on the strength of thinking so. That is the right order.
Curated picks where the model has the highest decision conviction. Updated every render.
Run rankings to populate Market Edge.
Run rankings first.
Mark a player as drafted to get a recommendation.
Studies and awards - the research-side surfaces of the managr platform.
Live snapshot of the projection system.
Backtest, parsimony, ablation, and benchmark - one synthesized report. Output renders in Validation.
Final Value = Raw Projection × MasterConfidence × TeamContext. Each layer answers one question - no overlap.
Recency weighting: prior seasons contribute by recency (more recent = more weight) and sample size (more PA = more weight).
Shrinkage: observed rates regress toward position-specific population means by an amount inversely proportional to sample size. Catches small-sample outliers.
Quality-of-contact adjustments: xwOBA, barrel rate, and bat-tracking metrics replace luck-driven outcome stats with skill-driven ones.
PT modeling: projected PA is anchored to workload tier (full-time, regular, platoon, backup catcher) using historical role distributions.
Monte Carlo: N simulated seasons per player using projected mean + uncertainty, producing P10 / P50 / P90 distribution and bust%.
Confidence layers: four orthogonal multipliers: Roster (depth chart), Role (PT certainty), Sample (career PA), Market (consensus alignment). The Audit tab shows all four for every player.
VAR: Value Above Replacement at the player's primary position. Replacement levels are floored per position so SS scarcity doesn't artificially inflate stable middle infielders.
Tune component weights against historical seasons. Stored in session.
-
-
-
Walk-forward backtests, calibration quality, and benchmark comparisons. Run walk-forward first - most tools depend on its pair pool.
Projects every historical batter-season from prior data only. Reports RMSE, MAE, Spearman ρ, hit rates. Populates the pair pool other validators use.
Walk-forward over historical seasons. Reports RMSE, MAE, Spearman ρ, top-N hit rates.
Distribution, uncertainty, per-archetype performance, calibrated tiers, draft sim.
Load any rival projection or ADP source. Auto-detects player_name + rank / adp / projected.
Where the model finds value the market is missing.
Curated picks from current rankings.
Where the model has measurable advantage.
The safety layer. Realism, false-confidence, forensics - what to remove.
The most important governance tools. Output renders in Validation.
Per-era ablation, error clusters, bias, correlations, removals. Each tells you what to remove. Run walk-forward first.
The smallest model that retains predictive power. Every feature beats the burden of proof.
Live pipeline status and data freshness.
-
Last known status of each underlying data source.
Source status will appear once data has been refreshed.
Issues from the most recent data refresh, if any.
-
MVP, Cy Young, Rookie of the Year - historical winners with their stats, plus model-predicted current-season winners.
Predictions from the current 2026 projection model. Top candidates per award based on projected fantasy points + advanced-stat underlying. League assignment from current team.
Coaching rosters for every MLB team plus a multi-year Fantasy Points Above Average ranking of every MLB coach.
Findings article + all-time top-100 performer leaderboard from every World Baseball Classic (2006-2026).
A growing collection of fantasy baseball studies. Each card opens a detail page with the methodology, the data sources, and (where the data exists today) live computed findings.
Active player injuries, refreshed from upstream sources.
| Player | Team | Pos | Status | Explanation | Replacement |
|---|
Enter players for each side, one per line. Values use your scoring weights and injury-adjusted projections.
| Player | Pos | ProjPts | Status |
|---|
| Player | Pos | ProjPts | Status |
|---|
Every signal, formula, and data source the model uses for player evaluation - organized by product.
Configure your draft session. This stays on this device only.