Where Sports Analytics Actually Beats Judgement

Statistical models beat unaided judgement on repetitive, well-measured questions. On single high-stakes calls, the published record is closer to a coin flip than anyone advertises.

By The Standard

September 9, 2026 4 min read· 812 words
Empty baseball dugout at dusk with a laptop and printed scouting reports on the bench, stadium floodlights on behind
In this articleObservation

The Standard take

Models beat unaided human judgement when the question is repetitive, well measured and high volume: season-long projections, draft-pick valuation, roster pricing. They do not beat it on single, low-sample, context-heavy calls — one draft pick, one athlete's injury timing — where the best published models sit uncomfortably close to a coin flip.

Evidence rating: High. This rating holds unless A pre-registered, blinded tournament pitting scouting boards against models on the same draft classes, with success metrics agreed in advance, would settle the draft question properly.

Key Findings

Evidence: High
  • Every winter someone declares that analytics has ruined sport, and every spring someone declares that anyone still trusting their eyes is a fossil.
  • Where do statistical models actually outperform experienced human judgement — and where does the published evidence say they do not?
  • Models beat unaided human judgement when the question is repetitive, well measured and high volume: season-long projections, draft-pick valuation, roster pricing.

What this does not prove: None of this shows models are superior for single high-stakes, low-sample decisions.

Every winter someone declares that analytics has ruined sport, and every spring someone declares that anyone still trusting their eyes is a fossil. Both camps are arguing about the same word as though it named one activity. It does not.

Where do statistical models actually outperform experienced human judgement — and where does the published evidence say they do not?

Models beat unaided human judgement when the question is repetitive, well measured and high volume: season-long projections, draft-pick valuation, roster pricing. They do not beat it on single, low-sample, context-heavy calls — one draft pick, one athlete's injury timing — where the best published models sit uncomfortably close to a coin flip.

The argument predates sport by fifty years. Paul Meehl's Clinical versus Statistical Prediction (1954) gathered roughly twenty head-to-head comparisons between expert clinicians and simple actuarial formulas, and the formulas won or tied nearly every time. The finding held up under scale: Grove, Zald, Lebow, Snitz and Nelson's meta-analysis of 136 studies found mechanical prediction about 10% more accurate on average than clinical judgement, and rarely worse.

That is the honest baseline. It is a real, replicated advantage. It is also 10%, not a revelation.

Draft valuation. Massey and Thaler's "The Loser's Curse" (Management Science, 2013) showed NFL decision-makers systematically overpay for top picks relative to the surplus value those picks return. The bias is predictable, quantifiable and exploitable by a model far simpler than the scouting apparatus it beats.

Projection systems. Here the humbling part. Tango's Marcel — three years of weighted stats, regressed to the mean, aged, and nothing else — performs close to commercial systems like PECOTA, ZiPS and Steamer in year-ahead accuracy tests. Annual accuracy audits find the gaps between systems measured in fractions of a win. The gain analytics delivered was over intuition, not over arithmetic.

Forecasting process. Tetlock's tournament work, and Mellers et al.'s superforecaster research, found that what separates good forecasters from bad is not expertise but process: aggregation, calibration, and updating. Trained amateurs beat credentialed analysts. That is a finding about method, not about machines.

Injury forecasting is the clearest failure case. Van Eetvelde and colleagues' systematic review of machine-learning injury prediction (Journal of Experimental Orthopaedics, 2021) found eleven qualifying studies with performance ranging from AUC 0.87 down to 0.52 — indistinguishable from a coin flip. Three of eleven were near chance. Most used small single-team samples with no external validation.

We work the same ground in Does training load predict injury risk?, where the methodological problem turns out to be worse than the sample-size problem.

Meehl's own literature contains the broken-leg case: the actuarial table predicts a man will attend the cinema tonight, the clinician knows he broke his leg this morning. When a human holds decisive information the model has no column for, the human should win — and in Grove's tally, clinical judgement did match or beat the model in a small but non-trivial minority of comparisons.

The second counterargument is the Marcel result read the other way. If a monkey-simple model matches an expensive one, the marginal value of modelling sophistication in this domain is small, and the industry's confidence in its own machinery is not fully earned.

A pre-registered, blinded tournament pitting scouting boards against models on the same draft classes, with success metrics agreed in advance, would settle the draft question properly. None exists. Single-cycle demonstrations that an algorithm "beat the scouts" are not that study. On the injury side, large multi-team datasets with genuine external validation would tell us whether the low AUCs reflect data scarcity or a genuinely unpredictable event.

None of this shows models are superior for single high-stakes, low-sample decisions. It shows they are superior on average, across many cases, where history is abundant. Those are different claims, and conflating them is how a front office talks itself into certainty it has not bought.

Based on the evidence presented, Second City Standard believes the useful question is not "models or scouts" but "how many times will this decision repeat?" Repetitive and well measured: use the model and stop arguing. Singular and context-heavy: use the model as a prior, and treat anyone claiming better than coin-flip accuracy on an individual injury as selling something.

More of how we handle this in Data & Intelligence, and the mechanics of one of these models in the WAR explainer.

Built By

Editorial Operating System v1.0

Creator

Erik Chambers

Architecture

Founder, Creator & Editorial Architect

AI Assisted

No

Human Reviewed

Pending

Evidence Reviewed

In progress

Last Updated

September 9, 2026

Confidence

Evergreen medium/10

Research Status

current

We do not claim perfection. We promise transparency. Every investigation shows its work — the question, the evidence, the tools, the humans, and the updates.

Analysis

Editorial Transparency

This article contains a combination of reporting, publicly available research, and editorial analysis.

Analysis and interpretation. Facts are sourced; conclusions are the author's. Evidence before opinion — facts require sources, analysis requires transparency, opinions require labels.

Meet the creator

Erik Chambers

Founder, Creator & Editorial Architect

Erik originated the central idea, directed the investigation, reviewed the evidence, and approved the final published work.

Read the founder profile →

Challenge This

We welcome disagreement

A different way to read the evidence. Research that points in another direction. Where specialists diverge.

Loading challenges…

Submit a challenge

Sign in to submit a challenge. All submissions are reviewed before appearing publicly.

Comments

Ask A Question

0

What would you ask an editor about this piece?

Reader questions feed our coverage map. The most-asked ones become our next investigations.

0/500

Continue Exploring

Guided by the evidence

Where should this take you next?

The Standard Score™

Where Sports Analytics Actually Beats Judgement

A composite 0–100 measure of how well this investigation meets the Second City Standard. Scores are auditable — every point comes from the criteria below.

Composite

of 100

Not yet scored

This investigation is queued for editorial scoring. No score has been assigned yet — the absence of a number is not a judgment on the evidence.

The Standard

· Editorial verdict

Based on the evidence presented,

Second City Standard believes Statistical models beat unaided judgement on repetitive, well-measured questions. On single high-stakes calls, the published record is closer to a coin flip than anyone advertises.

Remaining uncertainty: Awaiting a final written verdict from the editorial desk.

The Standard · Second City Standard

Continue The Investigation

· Never a dead end
The 2ND Take

Signal over noise.

One weekly dispatch. The best of 2ND CITY STANDARD, straight to your inbox.

No spam. Unsubscribe anytime.

You Set The Standard

Do you agree with this investigation?

Cast your verdict, join the discussion, or share it forward. Second City Standard grows when readers challenge the evidence.

Discuss