How accurate are sleep trackers, really?
The answer is more interesting than either camp wants it to be. These devices are considerably better than sceptics assume and considerably worse than the marketing implies — and the gap between the best and the median product in the category is enormous.
The best published consumer result is 79% agreement with clinical polysomnography on four-stage sleep classification. The benchmark that makes that number meaningful: two human technicians scoring the same night independently agree roughly 83% of the time.
But that is the ceiling, not the average. Across six wrist-worn devices tested independently, Cohen's kappa ran from 0.21 to 0.53 — fair to moderate agreement only. Sleep timing is reliable across the category. Sleep staging is not.
What "accurate" is being measured against
The reference standard is polysomnography: a night in a sleep laboratory with EEG electrodes on the scalp reading brain electrical activity, plus eye movement, muscle tone, airflow, respiratory effort and blood oxygen. A trained technician then scores the night in 30-second epochs into wake, N1, N2, N3 (deep) and REM.
The critical thing to understand before reading any accuracy figure: polysomnography itself is not perfect. Sleep staging involves human judgement, and two qualified technicians scoring the same recording independently agree only about 83% of the time. The "ground truth" has roughly 17% disagreement built into it.
Any device scoring above about 83% would not be more accurate than a sleep lab. It would be suspiciously well-fitted to one particular scorer.
The numbers
The headline result comes from a validation study conducted at Brigham and Women's Hospital, in which the Oura Ring reached 79% agreement with polysomnography on four-stage classification, with sensitivity of 76.0–79.5% and precision of 77.0–79.5% across stages. The peer-reviewed multi-sensor analysis found the ring did not differ significantly from polysomnography in its estimation of wake, light sleep, deep sleep or REM.
That is a genuinely strong result. It is also one device, in one study, on a particular population — and it is emphatically not representative of the category.
Cohen's kappa, and why it deflates the picture
Raw percentage agreement flatters any classifier, because some of that agreement happens by chance. If a device simply guessed "light sleep" for the entire night it would score respectably, because light sleep occupies most of a normal night. Cohen's kappa corrects for chance agreement, which makes it the honest metric.
By kappa, the picture is much more sober. A study of six commercial wrist-worn devices found values from 0.21 to 0.53 — conventionally read as fair to moderate. The strongest performers in that comparison were described as usable for tracking prolonged, significant changes in sleep architecture, which is a carefully limited claim: good enough to see a real shift over weeks, not good enough to trust last Tuesday's REM figure.
What they get right, and it matters more than the rest
Sleep timing is reliable. Sleep staging is not. Total sleep time, sleep onset and wake time are measured well by most modern devices — the underlying signals of movement and heart rate change decisively at those transitions. Distinguishing N2 from N3 from REM requires reading the brain, and no consumer device does.
This is not a consolation prize. For most people, sleep timing is the metric with actual consequences. Discovering that your true average is 6h10m rather than the 7h30m you assumed is a finding that changes behaviour. Whether 84 or 97 minutes of that was deep sleep changes nothing you can act on.
| Metric | Reliability | How to use it |
|---|---|---|
| Total sleep time | Good | Trust it. This is the number worth acting on. |
| Sleep onset / wake time | Good | Trust it. Useful for schedule consistency. |
| Resting heart rate | Good | Trust the trend. Reliable and genuinely informative. |
| Respiratory rate | Moderate | Trend only. A sustained rise is worth noting. |
| HRV | Moderate | Trend only, never a single night. See HRV explained. |
| Wake after sleep onset | Moderate | Devices generally under-detect brief awakenings. |
| Deep sleep | Weak | Estimate. Watch multi-week direction, not nightly values. |
| REM sleep | Weak | Estimate. Same caveat. |
| Composite "sleep score" | Not a measurement | A proprietary weighting of the above, unvalidated and not comparable between brands. |
Why rings tend to outperform watches
There is a physical reason, and it is not marketing. The finger has thinner tissue and denser capillary beds than the wrist, which yields a cleaner photoplethysmography signal — the optical pulse reading everything else is derived from. A ring also maintains constant contact, where a watch slides and rotates during the night, introducing motion artefacts precisely when the signal needs to be clean.
That said, form factor is not destiny: in the six-device wrist comparison, the best-performing wrist device reached kappa 0.53, ahead of several others by a wide margin. Implementation matters as much as placement. The practical trade-offs are covered in the smart ring buyer's guide and Oura vs Whoop.
Reading validation claims critically
Three questions to ask of any accuracy claim, including the ones on this page:
- Who funded and ran it? Manufacturer-run validation is not worthless — companies have the best access to their own hardware — but independent replication carries more weight. Note which you are reading.
- Which population? Studies in healthy young adults with normal sleep do not generalise to people with insomnia, sleep apnoea or shift work. Devices typically perform worse in disrupted sleep, which is exactly the population most motivated to buy one.
- Two-stage or four-stage? "Sleep versus wake" classification is a far easier problem than four-stage classification, and accuracy figures for it are much higher. A headline number without this specified is not comparable to anything.
There is a recognised pattern in which anxiety about tracked sleep quality actively degrades sleep. People lie awake concerned about their deep sleep percentage, or wake feeling fine and then feel worse after reading a poor score. Given that stage estimates are the least reliable numbers a device produces, being made anxious by them is a particularly poor trade. If your sleep score changes how your morning feels, that is a reason to stop looking at it daily.
Common questions
How accurate are sleep trackers compared to a sleep lab?
The best published consumer figure is 79% agreement with polysomnography on four-stage classification, from a Brigham and Women's Hospital study of the Oura Ring — against roughly 83% agreement between two human technicians scoring the same night. Across six wrist devices, Cohen's kappa ranged 0.21–0.53, only fair to moderate. The category ceiling is impressive; the category average is not.
Do sleep trackers measure brain waves?
No consumer wearable does. Polysomnography reads brain electrical activity directly via scalp EEG. Rings and watches infer sleep stages from motion, pulse waveform, heart rate variability, respiration and skin temperature — proxies for what the brain is doing, not measurements of it.
What are they most accurate at?
Sleep timing: total sleep time, onset and wake time. Most modern devices handle these reliably, and for the majority of people they are the metrics that actually change behaviour. Stage classification — particularly deep sleep and REM — should be treated as an estimate.
Can a tracker diagnose sleep apnoea?
No. It can flag patterns worth investigating — blood oxygen dips, elevated respiratory rate, frequent unexplained awakenings — and that prompt has real value; it has led to many legitimate referrals. But diagnosis requires a proper sleep study. Take the data to a clinician, not the conclusion.
Is a sleep score meaningful?
Less than it appears. A sleep score is a proprietary weighting of the underlying metrics, including the unreliable ones, and the formulas are neither published nor validated. Scores are not comparable between brands. The individual metrics are more informative than the number that summarises them.
Are these devices worth buying, given all this?
For many people, yes — but for the reliable metrics rather than the marketed ones. If you want to know how long you actually sleep, how consistent your schedule is, and whether an intervention like changing your bedroom temperature moved anything, a tracker does that well. If you want a trustworthy nightly REM figure, no consumer device delivers it.
Sources
- Four-stage agreement and inter-scorer benchmark — Brigham and Women's Hospital validation study summary (note: published by the manufacturer); peer-reviewed multi-sensor analysis at PMC8271886.
- Six-device wrist wearable comparison, Cohen's kappa 0.21–0.53 — PMC12038347.
- Three-device accuracy comparison — Sensors, MDPI 24(20):6532.
- Validation framework and methodology — PMC8161815.