Your Sleep Tracker Knows Less Than It Pretends To
Your Apple Watch says you got 1 hour 47 minutes of REM sleep last night. Here's what it actually measured, and why that's a very different thing.
By Erik Chambers
Founder, Creator & Editorial Architect

Built By
Editorial Operating System v1.0Creator
Erik Chambers
Architecture
Founder, Creator & Editorial Architect
AI Assisted
No
Human Reviewed
Pending
Evidence Reviewed
In progress
Last Updated
August 26, 2026
Confidence
Evergreen medium/10
Research Status
Living investigation
We do not claim perfection. We promise transparency. Every investigation shows its work — the question, the evidence, the tools, the humans, and the updates.
In this articleThe Question
Somewhere in Chicago this morning, someone glanced at their phone before they'd even opened their eyes fully and learned that they got 1 hour and 47 minutes of REM sleep, 58 minutes of deep sleep, and a "sleep score" of 82. The number arrived with the crisp, unblinking confidence of a lab report. No error bars. No asterisk. Just a clean pie chart, as if a piece of aluminum and glass on your wrist had spent the night reading your brain waves like a paperback.
It hadn't. That's not a knock on the technology — modern wearables are a genuine engineering achievement, packing an accelerometer, an optical heart-rate sensor, and sometimes a thermometer into something the size of a large button. But "impressive sensor package" and "clinical-grade sleep stage measurement" are different claims, and the gap between them is exactly where your sleep tracker's confident little pie chart starts writing checks its hardware can't cash.
This is the story of what's actually happening inside that device while you sleep, what it can genuinely tell you, and where the science says to raise an eyebrow at the numbers before you rearrange your evening routine around them.
The Question
When a consumer wearable reports precise sleep-stage minutes — REM, deep, light — how much of that is real measurement, and how much is an algorithm's best guess dressed up as data?
What We Know
None of the popular consumer wearables — not the Apple Watch, not Fitbit, not Oura — directly measure brain activity. The gold-standard clinical tool for sleep staging, polysomnography, works by attaching electrodes to the scalp to record electroencephalography (EEG), alongside eye movement (electrooculography) and muscle tone (electromyography) sensors, in an accredited sleep lab. That combination is what actually distinguishes REM sleep, light sleep, and deep slow-wave sleep according to established clinical scoring rules from the American Academy of Sleep Medicine.
Consumer wearables have none of that. What they actually have is an accelerometer that detects movement and stillness, a photoplethysmography (PPG) sensor that shines light into the skin to estimate heart rate and derive heart rate variability from the pattern of blood flow, and in some newer models, a skin temperature sensor and a pulse oximeter estimating blood oxygen saturation. From that bundle of movement, heart rate, and HRV data, a proprietary algorithm — built by Apple, Fitbit/Google, or Oura, and generally not published in full detail — makes a probabilistic guess about which sleep stage you're likely in during each stretch of the night, usually broken into short chunks called epochs.
Typical PSG scoring epoch
30seconds
Clinical sleep stages are scored in 30-second chunks from EEG; wearables approximate this using movement and heart signals instead of brain waves.
That inference is not baseless — there are genuine physiological correlations between sleep stage and heart rate variability (REM sleep tends to show more heart rate variability and irregular breathing patterns; deep sleep tends to show slower, steadier heart rate and minimal movement). The algorithms are exploiting real signal, not pulling numbers from thin air. But correlation-based inference from a different signal is categorically different from direct measurement, and that difference is precisely where accuracy degrades.
What the Data Says
Independent validation studies — researchers putting wearables and polysomnography on the same sleeping subjects on the same night and comparing results — have accumulated over the past several years across Apple Watch, Fitbit, and Oura devices, often published in venues like the Journal of Clinical Sleep Medicine and Sleep. The consistent pattern across this body of work: wearables tend to do reasonably well at the basic binary task of detecting sleep versus wake, and at estimating total sleep time in aggregate. They tend to be considerably weaker — and more inconsistent between devices and studies — at multi-stage classification, meaning correctly sorting a given epoch into REM, light, or deep sleep specifically.
| Task | General finding | Confidence level |
|---|---|---|
| Sleep vs. wake detection | Reasonably strong agreement with PSG in healthy adults | Moderate-to-high |
| Total sleep time estimate | Often close to PSG on average, though with individual-night error | Moderate |
| Detecting wake after sleep onset | Frequently underestimated (device sees "still" as "asleep") | Weak-to-moderate |
| REM vs. non-REM staging | Meaningful disagreement rate with PSG epoch-by-epoch | Weak |
| Light vs. deep sleep staging | Least reliable category, varies substantially by device | Weak |
Source: Synthesized from published validation studies comparing consumer wearables to polysomnography
One recurring, unglamorous finding across these studies is that wearables tend to have trouble specifically with "wake after sleep onset" — the periods when you're actually lying still but awake in the middle of the night. Because the devices lean heavily on stillness as a proxy for sleep, a person lying motionless but awake, scrolling their thoughts at 3 a.m., often gets misclassified as asleep. This is a specificity problem: the device is better at correctly flagging real sleep than it is at correctly flagging real wakefulness, which quietly inflates total sleep time and sleep efficiency numbers in a systematic direction.
The device is good at telling you that you slept. It is much less good at telling you what your brain was doing while you did.
Epoch-by-epoch agreement — comparing each 30-second window of wearable output against the simultaneous PSG scoring for that same window — is the most rigorous way researchers evaluate these devices, and it's a tougher test than comparing total nightly summaries, because two devices can arrive at similar total REM minutes for the night while disagreeing about which specific minutes were REM. Studies using this stricter epoch-level comparison generally find that agreement, while better than chance and improving with newer device generations, still falls meaningfully short of the reliability clinicians expect from PSG-to-PSG scorer comparisons.
Where the Evidence Gets Messy
A big complication: every major manufacturer treats its staging algorithm as proprietary and doesn't publish the full model. That means independent researchers are validating a black box, and when a company updates its algorithm — which happens periodically via software update — the accuracy profile can shift without a new public validation study necessarily following along at the same pace. A device you validated in 2021 is not guaranteed to be the same device, statistically, in 2024.
Validation populations also matter more than marketing tends to acknowledge. Many published studies use young, healthy adult volunteers in relatively controlled conditions, which is a reasonably favorable scenario for these algorithms. Accuracy for populations with sleep disorders, older adults, people with irregular schedules, or people with certain skin tones and PPG sensor performance differences has been less thoroughly studied across the board, and where it has been studied, performance sometimes drops — an important equity and reliability gap in the literature.
There's also a subtler statistical issue: manufacturers understandably promote their best validation numbers, often from studies they funded or co-authored, while independent, arm's-length replications are fewer and more scattered across different device models, firmware versions, and comparison methods, making it hard to build one clean, consensus accuracy figure that applies to "sleep trackers" as a category.
Second City Analysis
The Verdict
The claim that consumer sleep trackers reliably measure sleep stages is mostly supported for the coarse sleep-versus-wake and total-duration estimates, but overstated for the precise stage-by-stage numbers most apps display with equal confidence — the underlying validation literature shows real, non-trivial disagreement with polysomnography specifically at the staging level.
- Evidence strength
- 62
- Source quality
- 75
- Replication
- 55
- Sample quality
- 50
- Causation
- 20
- Scientific consensus
- 58
- Uncertainty
- 60
Sleep/wake and total-duration accuracy is reasonably well replicated across studies; stage-level (REM/deep) accuracy is inconsistent across devices, algorithm versions, and study populations, and proprietary algorithms limit independent scrutiny.
- 1.Accuracy of wearable devices for sleep tracking, Journal of Clinical Sleep Medicine — Link
- 2.Polysomnography scoring manual and standards, American Academy of Sleep Medicine — Link
- 3.Validation of consumer sleep-tracking technology against polysomnography, Sleep journal — Link
- 4.Photoplethysmography (PPG) sensing overview, National Institute of Biomedical Imaging and Bioengineering — Link
- 5.Consumer sleep technology position statement, American Academy of Sleep Medicine — Link
- 6.Orthosomnia: perfectionism and sleep-tracking anxiety, Journal of Clinical Sleep Medicine — Link
- 7.Heart rate variability and sleep stage physiology, National Institutes of Health / PubMed Central — Link
How We Measured This
- Question investigated
- How accurate are consumer sleep-tracking wearables, particularly their sleep-stage estimates, compared to clinical polysomnography?
- Evidence considered
- Published validation studies comparing wearable output (Apple Watch, Fitbit, Oura, and similar devices) against simultaneous polysomnography recordings, plus AASM clinical scoring standards and professional position statements on consumer sleep technology.
- Sources prioritised
- Peer-reviewed validation research in sleep medicine journals, AASM standards and position statements, and physiology literature on heart rate variability during sleep stages.
- Known limitations
- Proprietary algorithms limit full transparency; many validation studies use young healthy adults, limiting generalizability; device firmware and algorithms change over time faster than independent validation research can track.
- How the verdict was set
- Rated MOSTLY SUPPORTED for basic sleep/wake and duration detection, but the specific claim of precise stage-level accuracy is not well supported by the epoch-by-epoch validation literature, which shows meaningful disagreement with the polysomnography reference standard.
Editorial Transparency
This article contains a combination of reporting, publicly available research, and editorial analysis.
A long-form investigation. Findings resolve to primary sources. Evidence before opinion — facts require sources, analysis requires transparency, opinions require labels.
Meet the creator
Erik Chambers
Founder, Creator & Editorial Architect
Erik originated the central idea, directed the investigation, reviewed the evidence, and approved the final published work.
Read the founder profile →Challenge This
We welcome disagreement
A different way to read the evidence. Research that points in another direction. Where specialists diverge.
Loading challenges…
Submit a challenge
Sign in to submit a challenge. All submissions are reviewed before appearing publicly.
Comments
Ask A Question
What would you ask an editor about this piece?
Reader questions feed our coverage map. The most-asked ones become our next investigations.
Continue Exploring
Guided by the evidenceWhere should this take you next?
Your Sleep Tracker Knows Less Than It Pretends To
A composite 0–100 measure of how well this investigation meets the Second City Standard. Scores are auditable — every point comes from the criteria below.
Composite
—
of 100
This investigation is queued for editorial scoring. No score has been assigned yet — the absence of a number is not a judgment on the evidence.
The Standard
· Editorial verdictBased on the evidence presented,
Second City Standard believes Your Apple Watch says you got 1 hour 47 minutes of REM sleep last night. Here's what it actually measured, and why that's a very different thing.
Remaining uncertainty: Awaiting a final written verdict from the editorial desk.
The Standard · Second City Standard
Continue The Investigation
· Never a dead endThis investigation is one thread. Pull the next one — every path below is another investigation, another question, or another discipline applied to the same problem.
Signal over noise.
One weekly dispatch. The best of 2ND CITY STANDARD, straight to your inbox.
No spam. Unsubscribe anytime.