What a Win Is Actually Worth
Every contract in baseball is priced off a number that two respected sources cannot agree on. That disagreement is not a scandal — it is the whole story.
By The Standard
In this articleObservation
Key Findings
- Two front offices can look at the same player, run the same season, and arrive at different valuations — not because one is lazy, but because the public metric they are both nominally using does not have one definition.
- If WAR is the number that prices players, what exactly is it measuring, and how much confidence does a single WAR figure actually deserve?
- Before WAR, value arguments ran on counting stats: home runs, RBI, pitcher wins.
Observation
Two front offices can look at the same player, run the same season, and arrive at different valuations — not because one is lazy, but because the public metric they are both nominally using does not have one definition.
Wins Above Replacement is the closest thing baseball has to a single-number verdict. It appears in arbitration filings, in television graphics, in Hall of Fame arguments, and in the sentence "he was worth X million." Yet the two best-known implementations — FanGraphs (fWAR) and Baseball-Reference (bWAR) — regularly disagree about the same pitcher by a full win or more.
The Question
If WAR is the number that prices players, what exactly is it measuring, and how much confidence does a single WAR figure actually deserve?
Historical Context
Before WAR, value arguments ran on counting stats: home runs, RBI, pitcher wins. Those numbers reward opportunity as much as ability — RBI depends on who bats in front of you, pitcher wins depend on who bats behind you. Sabermetric work through the 1980s and 1990s, popularised by Bill James and later absorbed into front offices, replaced "what happened while he was there" with "how much did he change the run environment."
WAR is the end point of that logic: convert everything a player does into runs, convert runs into wins, then subtract what a freely available replacement-level player would have produced in the same playing time.
Scientific Background
The chain is: events → runs → wins → value above replacement.
- Events to runs uses linear weights — each outcome (walk, single, home run, out) gets a run value based on how much it changes expected run scoring.
- Runs to wins uses a run-per-win converter that varies with the scoring environment; roughly ten runs to a win in a typical modern context, fewer in a low-scoring era.
- Replacement level is a defined baseline, not a measured one. It is a convention chosen so that the sum of all player value lands near the league's actual win total.
Every one of those steps is a modelling decision. That is where the two WARs part company.
Data and Metrics
The single largest source of divergence is pitcher evaluation.
- FanGraphs' fWAR prices pitchers primarily on FIP — strikeouts, walks, hit batters, home runs — the outcomes least dependent on the defenders behind them.
- Baseball-Reference's bWAR prices pitchers on runs actually allowed, then adjusts for the quality of the defense behind them and the parks they threw in.
A pitcher who strikes out few hitters but posts a low ERA in front of an excellent defense will look mediocre by fWAR and excellent by bWAR. Neither system is malfunctioning. They are answering different questions: what did this pitcher control? versus what actually happened while he pitched?
Position players carry a second uncertainty: defensive metrics. Batting value is estimated from thousands of plate appearances and is comparatively stable. Fielding value is estimated from far fewer meaningful chances, and different fielding systems can disagree about the same shortstop by a win.
Park effects add a third. A neutral-park hitter and a Coors Field hitter with identical raw lines are not identical hitters, and the size of that correction depends on how many seasons of park data a system smooths over.
Human Factors
Front offices know all of this. Agents know it too, and cite whichever version is kinder. Arbitration panels hear both. The public conversation, meanwhile, usually cites one number without a source, which is how a modelling estimate acquires the tone of a measurement.
The result is a market that behaves as if WAR were precise while the people trading on it treat it as a range.
Counterarguments
"The two systems agree most of the time." For everyday position players over full seasons, largely true — the disagreements concentrate in pitchers, part-time players, and defense-dependent profiles. That is precisely where expensive mistakes live.
"Teams have better internal metrics." Almost certainly. Clubs have tracking data the public does not. But better inputs do not remove the structural choices — replacement level is still a convention, and runs-to-wins is still a model.
"Dollars per win fixes it." Dollars per win is an inference drawn from observed free-agent contracts. It describes what the market paid, not what a win is worth. It also mixes in factors WAR never claimed to price: age, marketing value, roster fit, deferred money, and the club's own competitive window.
Evidence
What the record supports:
- WAR is an estimate assembled from several models, each with its own error, not a measurement.
- The two public versions differ most where the underlying data is thinnest — pitching context and defense.
- Market prices for wins are derived from contracts, so using them to "value" a contract is partly circular.
What the record does not support: any claim that a single decimal-place WAR difference between two players establishes which is better.
Conclusions
A win has a price, but the price is a negotiated summary of several uncertain estimates. The honest way to read WAR is as a range with a centre, not a verdict — a five-win player is a player whose most likely value is around five wins, with real error bars in both directions.
Remaining Questions
- How much of the fWAR/bWAR gap for pitchers closes as batted-ball tracking improves defensive attribution?
- Does the dollars-per-win figure hold across contract lengths, or does it mostly describe short deals for players in their late twenties?
- Where should a club's internal replacement level sit when its own depth is unusually strong or weak?
The Standard
We treat WAR as evidence, not as a scoreboard. Where a valuation depends on which version of WAR is used, we say so, and we show both. Where the difference between two players is smaller than the disagreement between the systems measuring them, the correct answer is that we do not know.
Built By
Editorial Operating System v1.0Creator
Erik Chambers
Architecture
Founder, Creator & Editorial Architect
AI Assisted
No
Human Reviewed
Pending
Evidence Reviewed
In progress
Last Updated
September 1, 2026
Confidence
Evergreen high/10
Research Status
Living investigation
We do not claim perfection. We promise transparency. Every investigation shows its work — the question, the evidence, the tools, the humans, and the updates.
Editorial Transparency
This article contains a combination of reporting, publicly available research, and editorial analysis.
Analysis and interpretation. Facts are sourced; conclusions are the author's. Evidence before opinion — facts require sources, analysis requires transparency, opinions require labels.
Meet the creator
Erik Chambers
Founder, Creator & Editorial Architect
Erik originated the central idea, directed the investigation, reviewed the evidence, and approved the final published work.
Read the founder profile →Challenge This
We welcome disagreement
A different way to read the evidence. Research that points in another direction. Where specialists diverge.
Loading challenges…
Submit a challenge
Sign in to submit a challenge. All submissions are reviewed before appearing publicly.
Comments
Ask A Question
What would you ask an editor about this piece?
Reader questions feed our coverage map. The most-asked ones become our next investigations.
Continue Exploring
Guided by the evidenceWhere should this take you next?
Part of a cluster
Baseball Value
What is a baseball player actually worth, and who decides?
See the whole clusterWhat remains open
- How much of the fWAR/bWAR gap for pitchers closes as batted-ball tracking improves defensive attribution?
- Does the dollars-per-win figure hold across contract lengths, or does it mostly describe short deals for players in their late twenties?
- Where should a club's internal replacement level sit when its own depth is unusually strong or weak?
What a Win Is Actually Worth
A composite 0–100 measure of how well this investigation meets the Second City Standard. Scores are auditable — every point comes from the criteria below.
Composite
—
of 100
This investigation is queued for editorial scoring. No score has been assigned yet — the absence of a number is not a judgment on the evidence.
The Standard
· Editorial verdictBased on the evidence presented,
Second City Standard believes Every contract in baseball is priced off a number that two respected sources cannot agree on. That disagreement is not a scandal — it is the whole story.
Remaining uncertainty: Awaiting a final written verdict from the editorial desk.
The Standard · Second City Standard
Continue The Investigation
· Never a dead endThis investigation is one thread. Pull the next one — every path below is another investigation, another question, or another discipline applied to the same problem.
Signal over noise.
One weekly dispatch. The best of 2ND CITY STANDARD, straight to your inbox.
No spam. Unsubscribe anytime.