Consumer sleep trackers are better at the simple question than the interesting one. Asked whether you were asleep, they are broadly reliable. Asked which stage of sleep you were in, they are considerably less so, and that is the number most people actually look at in the morning.

The clearest evidence comes from a study of seven consumer sleep-tracking devices compared with polysomnography, published in the journal SLEEP in 2021 by Chinoy and colleagues. Polysomnography is the clinical standard, measuring brain activity directly. The devices were tested against it in 34 healthy young adults across three consecutive nights. This piece is part of our complete guide to better sleep.

How trackers estimate sleep in the first place

This is the root of everything below. Clinical sleep staging reads brain activity. A wrist device does not have access to that, so it infers sleep from what it can measure, typically movement and heart rate, and then applies a manufacturer's algorithm to turn those signals into a hypnogram of light, deep and REM sleep.

That inference is reasonable for the coarse question. Lying still with a slow, steady heart rate is a decent proxy for being asleep. It is a much weaker basis for deciding which stage of sleep you were in, and it is a poor basis for spotting the minutes you spent awake while lying still.

Where they do well, and where they do not

The study's conclusion on the basics is genuinely positive: "Consumer sleep-tracking devices exhibited high performance in detecting sleep, and most performed equivalent to (or better than) actigraphy in detecting wake." Actigraphy is the research-grade movement monitor these devices are usually benchmarked against, so matching or beating it is a real result.

The detail underneath that is where the nuance sits: "epoch-by-epoch sensitivity was high (all ≥0.93), specificity was low-to-medium (0.18–0.54)."

Measure What it means here Result
Sensitivity Correctly calling sleep when you were asleep High, all devices at 0.93 or above
Specificity Correctly calling wake when you were awake Low to medium, 0.18 to 0.54
Sleep staging Distinguishing light, deep and REM Inconsistent across devices

Read the second row carefully, because it is the practical catch. A device that is excellent at recognising sleep and weak at recognising wakefulness will tend to count quiet wakefulness as sleep. The 2017 Journal of Clinical Sleep Medicine paper discussed below makes the same point plainly, noting that these devices "tend to overestimate sleep" and have "poor accuracy in detecting wake after sleep onset."

On staging, the study's verdict was that "device sleep stage assessments were inconsistent." That is the sentence to remember the next time an app tells you that you got 47 minutes of deep sleep.

Two limits on all of this are worth stating. The study tested named consumer models available at the time, and its own authors recommended that devices "should be tested in different populations and settings to further examine their wider validity and utility." We are not going to rank products or tell you which to buy: the specific models tested are years old now, and a study in 34 healthy young adults is not a verdict on every device or every sleeper.

The risk of taking the data too seriously

There is a documented failure mode here, and it has a name. In Orthosomnia: Are Some Patients Taking the Quantified Self Too Far?, Baron and colleagues describe patients who are "preoccupied or concerned with improving or perfecting their wearable sleep data." They built the term from "ortho" meaning straight or correct and "somnia" meaning sleep, and they draw the parallel explicitly: "the perfectionist quest to achieve perfect sleep is similar to the unhealthy preoccupation with healthy eating, termed orthorexia."

The mechanism they describe is uncomfortably familiar. "The patients' inferred correlation between sleep tracker data and daytime fatigue may become a perfectionistic quest for the ideal sleep." You feel tired, the app offers a number to explain it, and the number becomes the target.

What makes this more than a psychological curiosity is that the same authors are blunt about the data driving it: consumer devices "are unable to accurately discriminate stages of sleep and have poor accuracy in detecting wake after sleep onset," and "lack of transparency in the device algorithms makes it impossible to know how accurate they are." So the anxiety is often being generated by the least reliable figure on the screen.

How to use one usefully

  1. Trust the trend, not the night. Weeks of bedtimes and rough durations tell you something. Tuesday's sleep score does not.
  2. Use it for timing and consistency. The most useful thing a tracker records is when you went to bed and when you got up, which maps directly onto the habit CDC/NIOSH leads with: the same times every day, including days off.
  3. Discount the stage breakdown. Deep and REM percentages are the part the research found inconsistent, and no amount of staring changes what you did overnight.
  4. Judge the day, not the dashboard. The CDC's own markers of poor sleep are things you notice awake: trouble falling asleep, waking repeatedly, and feeling tired despite enough sleep. Those outrank any score.
  5. If checking it makes you anxious, stop checking it. That is not a failure of discipline. It is the specific pattern the orthosomnia paper describes.

If a tracker is telling you something is wrong and your days agree, the useful next step is not a better device. It is the everyday causes in why you cannot fall asleep, the habits in our sleep hygiene checklist, or a clinician if the pattern persists.