How we test
Every device review on this site follows the same protocol. It is published here in full so you can decide how much weight the conclusions deserve — and so you can point out where it is wrong.
Devices are tested for a minimum of 30 consecutive nights after a 14-night unmonitored baseline. Where two devices make overlapping claims, they are worn simultaneously for at least 14 of those nights. Nightly exports are kept as raw CSV and the summary statistics are published with the review. We buy the hardware at retail wherever the budget allows; where a unit is supplied by a manufacturer, the review says so in the first screen of text.
Why 30 nights
Because shorter tests measure novelty, not the device. Sleep varies enormously night to night for reasons that have nothing to do with what is on your finger — alcohol, room temperature, illness, a late meal, an argument. A week of data is dominated by that noise. Thirty nights is the point at which a stable per-device mean starts to emerge for the metrics these products report, and it is long enough to catch the failure modes that only appear with use: strap irritation, charge-cycle degradation, sync failures, an app update that silently changes an algorithm.
It is also long enough to get past the first-week effect, where simply knowing you are being measured changes your bedtime.
The protocol, step by step
Baseline
Two weeks of normal sleep with no new device, no schedule change, and no deliberate optimisation. Bedroom temperature, bed and bedding are held constant from this point until the test ends.
Acclimatisation
The device is worn or installed and used exactly as a normal buyer would, following the manufacturer's own setup. Data from this week is collected but excluded from the published averages — this is the week where you are still adjusting the strap and being kept awake by a new gadget.
Measurement window
The 23 nights that count. Nightly export of every metric the device makes available, plus a fixed morning log: subjective rest 1–10, alcohol units, caffeine after 2pm, bedroom temperature, and any illness. These confounders are recorded so an odd week can be explained rather than quietly averaged away.
Head-to-head overlap
Where a competing device makes the same claim, both are worn on the same nights for at least 14 nights. Agreement is reported as the mean absolute difference per metric, not as a winner. Two devices disagreeing by 22 minutes on deep sleep tells you something useful about both.
Live-with period
Testing does not stop at publication. Reviews carry a "last verified" date, and pages are revisited when the manufacturer ships a significant firmware or algorithm change — which in this category happens more often than you would expect.
What we compare against
Nobody outside a sleep laboratory has ground truth. Clinical polysomnography — electrodes on the scalp, face and chest, scored by a technician — is the reference standard, and it is not something a review site can run. What we can do is be precise about the substitute.
| Measurement | Reference used | What it can and can't settle |
|---|---|---|
| Sleep timing onset, wake, total sleep time |
Manual sleep diary plus a second independent device | Reliable for total sleep time and wake time. A device that says you slept 7h40m when you were reading until 1am is straightforwardly wrong, and this catches it. |
| Sleep staging light, deep, REM |
Cross-device agreement only | Cannot be settled. Without EEG there is no truth to compare to. We report how far devices disagree with each other and defer to published validation studies for accuracy claims. |
| Heart rate & HRV | Chest-strap ECG monitor on overlapping nights | Good. A chest strap reads the electrical signal directly rather than inferring it from an optical pulse, so it is a fair check on nocturnal averages. |
| Bed and room temperature | Independent logging thermometer, surface and ambient | Good. This is the one place in the category where a cheap external sensor genuinely arbitrates — it catches cooling systems that do not reach their claimed set point. |
| Blood biomarkers | Accredited laboratory, same lab and same assay across draws | Good, with caveats. Assays differ between labs, so a level from one provider is not directly comparable to another's. We keep the lab constant within a comparison. |
No consumer device on this site measures brain activity. Sleep stages are inferred from motion, pulse waveform, heart-rate variability, respiration and skin temperature. When a review reports deep sleep, it is reporting the device's estimate, and any statement about accuracy comes from published validation research — not from our test. See how accurate are sleep trackers for what that research actually found.
How devices are bought
- Retail purchase is the default. Buying the product the way a reader buys it also tests the things manufacturers control when they know a reviewer is watching: delivery, sizing, returns, and what happens when you cancel a subscription.
- Loaned or supplied units are labelled. If a manufacturer provides hardware, it is stated in the first screen of the review, not in a footnote. Supplied units are returned or the equivalent value is donated at the end of testing.
- No manufacturer sees a review before publication. No copy approval, no fact-check window, no embargo trade in exchange for early access. If that costs access to a launch, we are late to that launch.
- Commission does not enter the ranking. Affiliate rates are checked only after conclusions are written. Several products recommended on this site pay nothing at all.
How we score
There is no composite score out of ten on this site. A single number implies a shared unit of measure between "battery life" and "how the app makes you feel," and there isn't one. Instead every review ends with the same three-part verdict:
- Who it is for — the specific buyer whose situation this device fits, stated narrowly enough to exclude people.
- Who should skip it — the buyer for whom it is the wrong purchase, and what they should look at instead.
- The five-year cost — hardware plus subscription plus the replacement cycle, because in this category the sticker price is routinely under half of what you actually spend.
Corrections
Errors are fixed on the page with a dated note at the foot of the article, not silently. If a conclusion changes because a device improved, the old conclusion stays visible alongside the new one. Testing notes and raw exports are available on request to anyone who wants to check the arithmetic — the email address is on the about page.
This protocol is published as version 1.0 on 6 September 2026. Reviews currently inside their measurement window are marked In testing and show no results until the window closes. That label is not a placeholder for content that exists elsewhere — it means the nights have not been slept yet.