Does WAR Actually Predict Team Wins?

Summed player value tracks a season's standings closely in aggregate and misses individual clubs by several games. The gap is the interesting part.

By The Standard

September 1, 2026 4 min read· 813 words
Empty baseball stadium at dusk seen from high behind home plate, floodlights coming on over an empty infield
The standings are settled on this field. The model is settled somewhere else.Photo: Second City Standard
In this articleObservation

Key Findings

  • Every winter, a front office adds up projected WAR for forty players, adds the replacement-level baseline, and produces a win total.
  • Does summed Wins Above Replacement predict team wins — and if it does, why does it keep missing specific teams by margins that decide playoff races?
  • Summed WAR plus a replacement baseline lands close to actual team wins in aggregate, and misses individual clubs by roughly four to six games in a typical season.

Every winter, a front office adds up projected WAR for forty players, adds the replacement-level baseline, and produces a win total. Every October, some of those teams finish six games off the projection in one direction or the other, and everyone acts surprised.

Both things are normal. They are, in fact, the same fact stated twice.

Does summed Wins Above Replacement predict team wins — and if it does, why does it keep missing specific teams by margins that decide playoff races?

Summed WAR plus a replacement baseline lands close to actual team wins in aggregate, and misses individual clubs by roughly four to six games in a typical season. The model is sound; the residual is real. Sequencing, bullpen leverage, injuries, and playing-time distribution live entirely outside the equation, and they are exactly what separates an 84-win season from a 90-win one.

WAR was built to answer a roster question, not a standings question. FanGraphs' primer on the framework is explicit that replacement level is a chosen convention: it is set so that a roster of freely available players would win somewhere in the neighbourhood of 47 to 48 games over a 162-game season, and so that total league WAR reconciles with total league wins above that floor. Baseball-Reference's version makes the same reconciliation with different pitching and defensive inputs.

That reconciliation is the whole trick. Because the baseline is calibrated to league totals, league-wide summed WAR has to line up with league-wide wins. Nothing in that construction promises it lines up for the Brewers.

Team wins, in WAR's language, are: replacement baseline, plus position-player WAR, plus pitching WAR. Each player's contribution is a run total divided by roughly ten runs per win.

Three things about that arithmetic matter more than they get credit for.

The components are added as if independent. Bat, glove, legs, and position are summed with no interaction term, even though a defensive alignment is a team property and a lineup's run scoring is not the sum of nine solo performances.

The two public versions disagree with each other on purpose. fWAR grades pitchers on FIP — strikeouts, walks, home runs — while bWAR grades them on runs actually allowed, adjusted for the defense behind them. Baseball-Reference publishes a direct comparison of the two frameworks. A rotation can be worth several wins more in one system than the other, which means "the team's WAR" is not one number.

Context is stripped deliberately. WAR does not know that the home run came with the bases loaded. That is a design decision that makes the metric more predictive going forward and less descriptive of what already happened — and team wins are, by definition, a description of what already happened.

The residual is not random noise scattered evenly. It clusters in identifiable places.

Sequencing. A team that hits with runners on base outperforms its component stats; one that does not, underperforms. Neither is in the model, and neither reliably repeats.

Bullpen leverage. Relievers pitch a small number of innings that swing win probability enormously. WAR applies a leverage adjustment, but the year's actual distribution of high-leverage outcomes is far coarser than any season-total adjustment can capture.

Playing time. Preseason WAR projections assume a distribution of plate appearances and innings. Injuries redistribute them to worse players. The projection was not wrong about the players; it was wrong about who would play.

Defensive uncertainty. The fielding term carries the widest error bar of anything in the equation — a point we work through in full in the WAR explainer. Aggregated across a roster, those errors partially cancel. Partially.

A version of team WAR that reconciled to individual clubs — not just the league — within two games in most seasons, using only preseason inputs, would mean the residual is model error rather than genuine in-season variance. Public tracking data on defensive positioning and reliever deployment is the most likely source of that improvement, and it would move this from "useful with a range" to "usable as a forecast."

Use summed WAR the way you would use a weather forecast: directionally right, honest about its range, useless as a promise. It tells you whether a roster is built to win 78 games or 92. It does not tell you which one it will win, and any front office that treats it as though it does has confused a model of talent for a model of outcomes.

Run your own numbers in the WAR Calculator, and see how the rest of our analytics work fits together in Data & Intelligence.

Built By

Editorial Operating System v1.0

Creator

Erik Chambers

Architecture

Founder, Creator & Editorial Architect

AI Assisted

No

Human Reviewed

Pending

Evidence Reviewed

In progress

Last Updated

September 1, 2026

Confidence

Evergreen medium/10

Research Status

Living investigation

We do not claim perfection. We promise transparency. Every investigation shows its work — the question, the evidence, the tools, the humans, and the updates.

Analysis

Editorial Transparency

This article contains a combination of reporting, publicly available research, and editorial analysis.

Analysis and interpretation. Facts are sourced; conclusions are the author's. Evidence before opinion — facts require sources, analysis requires transparency, opinions require labels.

Meet the creator

Erik Chambers

Founder, Creator & Editorial Architect

Erik originated the central idea, directed the investigation, reviewed the evidence, and approved the final published work.

Read the founder profile →

Challenge This

We welcome disagreement

A different way to read the evidence. Research that points in another direction. Where specialists diverge.

Loading challenges…

Submit a challenge

Sign in to submit a challenge. All submissions are reviewed before appearing publicly.

Comments

Ask A Question

0

What would you ask an editor about this piece?

Reader questions feed our coverage map. The most-asked ones become our next investigations.

0/500

Continue Exploring

Guided by the evidence

Where should this take you next?

Part of a cluster

Baseball Value

What is a baseball player actually worth, and who decides?

See the whole cluster
The Standard Score™

Does WAR Actually Predict Team Wins?

A composite 0–100 measure of how well this investigation meets the Second City Standard. Scores are auditable — every point comes from the criteria below.

Composite

of 100

Not yet scored

This investigation is queued for editorial scoring. No score has been assigned yet — the absence of a number is not a judgment on the evidence.

The Standard

· Editorial verdict

Based on the evidence presented,

Second City Standard believes Summed player value tracks a season's standings closely in aggregate and misses individual clubs by several games. The gap is the interesting part.

Remaining uncertainty: Awaiting a final written verdict from the editorial desk.

The Standard · Second City Standard

Continue The Investigation

· Never a dead end
The 2ND Take

Signal over noise.

One weekly dispatch. The best of 2ND CITY STANDARD, straight to your inbox.

No spam. Unsubscribe anytime.

You Set The Standard

Do you agree with this investigation?

Cast your verdict, join the discussion, or share it forward. Second City Standard grows when readers challenge the evidence.

Discuss