Does WAR Actually Predict Team Wins?
Summed player value tracks a season's standings closely in aggregate and misses individual clubs by several games. The gap is the interesting part.
By The Standard

In this articleObservation
Key Findings
- Every winter, a front office adds up projected WAR for forty players, adds the replacement-level baseline, and produces a win total.
- Does summed Wins Above Replacement predict team wins — and if it does, why does it keep missing specific teams by margins that decide playoff races?
- Summed WAR plus a replacement baseline lands close to actual team wins in aggregate, and misses individual clubs by roughly four to six games in a typical season.
Observation
Every winter, a front office adds up projected WAR for forty players, adds the replacement-level baseline, and produces a win total. Every October, some of those teams finish six games off the projection in one direction or the other, and everyone acts surprised.
Both things are normal. They are, in fact, the same fact stated twice.
The Question
Does summed Wins Above Replacement predict team wins — and if it does, why does it keep missing specific teams by margins that decide playoff races?
The Standard take
Summed WAR plus a replacement baseline lands close to actual team wins in aggregate, and misses individual clubs by roughly four to six games in a typical season. The model is sound; the residual is real. Sequencing, bullpen leverage, injuries, and playing-time distribution live entirely outside the equation, and they are exactly what separates an 84-win season from a 90-win one.
Historical Context
WAR was built to answer a roster question, not a standings question. FanGraphs' primer on the framework is explicit that replacement level is a chosen convention: it is set so that a roster of freely available players would win somewhere in the neighbourhood of 47 to 48 games over a 162-game season, and so that total league WAR reconciles with total league wins above that floor. Baseball-Reference's version makes the same reconciliation with different pitching and defensive inputs.
That reconciliation is the whole trick. Because the baseline is calibrated to league totals, league-wide summed WAR has to line up with league-wide wins. Nothing in that construction promises it lines up for the Brewers.
What the Model Actually Says
Team wins, in WAR's language, are: replacement baseline, plus position-player WAR, plus pitching WAR. Each player's contribution is a run total divided by roughly ten runs per win.
Three things about that arithmetic matter more than they get credit for.
The components are added as if independent. Bat, glove, legs, and position are summed with no interaction term, even though a defensive alignment is a team property and a lineup's run scoring is not the sum of nine solo performances.
The two public versions disagree with each other on purpose. fWAR grades pitchers on FIP — strikeouts, walks, home runs — while bWAR grades them on runs actually allowed, adjusted for the defense behind them. Baseball-Reference publishes a direct comparison of the two frameworks. A rotation can be worth several wins more in one system than the other, which means "the team's WAR" is not one number.
Context is stripped deliberately. WAR does not know that the home run came with the bases loaded. That is a design decision that makes the metric more predictive going forward and less descriptive of what already happened — and team wins are, by definition, a description of what already happened.
Where It Breaks
The residual is not random noise scattered evenly. It clusters in identifiable places.
Sequencing. A team that hits with runners on base outperforms its component stats; one that does not, underperforms. Neither is in the model, and neither reliably repeats.
Bullpen leverage. Relievers pitch a small number of innings that swing win probability enormously. WAR applies a leverage adjustment, but the year's actual distribution of high-leverage outcomes is far coarser than any season-total adjustment can capture.
Playing time. Preseason WAR projections assume a distribution of plate appearances and innings. Injuries redistribute them to worse players. The projection was not wrong about the players; it was wrong about who would play.
Defensive uncertainty. The fielding term carries the widest error bar of anything in the equation — a point we work through in full in the WAR explainer. Aggregated across a roster, those errors partially cancel. Partially.
What would change our mind
A version of team WAR that reconciled to individual clubs — not just the league — within two games in most seasons, using only preseason inputs, would mean the residual is model error rather than genuine in-season variance. Public tracking data on defensive positioning and reliever deployment is the most likely source of that improvement, and it would move this from "useful with a range" to "usable as a forecast."
The Standard
Use summed WAR the way you would use a weather forecast: directionally right, honest about its range, useless as a promise. It tells you whether a roster is built to win 78 games or 92. It does not tell you which one it will win, and any front office that treats it as though it does has confused a model of talent for a model of outcomes.
Run your own numbers in the WAR Calculator, and see how the rest of our analytics work fits together in Data & Intelligence.
Built By
Editorial Operating System v1.0Creator
Erik Chambers
Architecture
Founder, Creator & Editorial Architect
AI Assisted
No
Human Reviewed
Pending
Evidence Reviewed
In progress
Last Updated
September 1, 2026
Confidence
Evergreen medium/10
Research Status
Living investigation
We do not claim perfection. We promise transparency. Every investigation shows its work — the question, the evidence, the tools, the humans, and the updates.
Editorial Transparency
This article contains a combination of reporting, publicly available research, and editorial analysis.
Analysis and interpretation. Facts are sourced; conclusions are the author's. Evidence before opinion — facts require sources, analysis requires transparency, opinions require labels.
Meet the creator
Erik Chambers
Founder, Creator & Editorial Architect
Erik originated the central idea, directed the investigation, reviewed the evidence, and approved the final published work.
Read the founder profile →Challenge This
We welcome disagreement
A different way to read the evidence. Research that points in another direction. Where specialists diverge.
Loading challenges…
Submit a challenge
Sign in to submit a challenge. All submissions are reviewed before appearing publicly.
Comments
Ask A Question
What would you ask an editor about this piece?
Reader questions feed our coverage map. The most-asked ones become our next investigations.
Continue Exploring
Guided by the evidenceWhere should this take you next?
Part of a cluster
Baseball Value
What is a baseball player actually worth, and who decides?
See the whole clusterDoes WAR Actually Predict Team Wins?
A composite 0–100 measure of how well this investigation meets the Second City Standard. Scores are auditable — every point comes from the criteria below.
Composite
—
of 100
This investigation is queued for editorial scoring. No score has been assigned yet — the absence of a number is not a judgment on the evidence.
The Standard
· Editorial verdictBased on the evidence presented,
Second City Standard believes Summed player value tracks a season's standings closely in aggregate and misses individual clubs by several games. The gap is the interesting part.
Remaining uncertainty: Awaiting a final written verdict from the editorial desk.
The Standard · Second City Standard
Continue The Investigation
· Never a dead endThis investigation is one thread. Pull the next one — every path below is another investigation, another question, or another discipline applied to the same problem.
Signal over noise.
One weekly dispatch. The best of 2ND CITY STANDARD, straight to your inbox.
No spam. Unsubscribe anytime.