Sentinel / Legal Research · Claim Assessment

Oura “95% Sleep Staging Accuracy”

A proposed class action filed in the U.S. District Court for the Northern District of California by plaintiff Madison Surber, represented by Clarkson Law Firm, challenges Oura’s public “95% Sleep Staging Accuracy Compared to clinical sleep lab” representation. The lawsuit alleges that Oura overstates its ability to identify sleep stages using a finger-worn device. This assessment does not originate that claim. It analyzes the challenged representation against identified public evidence, including Oura-affiliated explanations, independent studies, and the specific measurement task attached to each reported percentage.

The advertised number does not match the evidence for the advertised task.

From the evidence analyzed and traced to the specific supporting artifacts, the 95% figure is associated with two-class sleep/wake classification, not the four-stage sleep-staging task communicated by the challenged claim. The traced Oura-affiliated evidence identifies 79% for four-stage classification, while an independent clinical-context study reports 53.18% under its tested population and protocol. That artifact-to-claim comparison produced the conclusion above.

The system produces the finding. The explanation layer makes the finding readable. Sentinel reports what the identified record supports, contradicts, limits, or does not establish.
01 · Scope & boundary

What question is being assessed?

What does the presently identified evidence show about the relationship among the advertised 95% sleep-staging statement, the 95% two-class sleep/wake figure, the 79% four-class figure, and the independent 53.18% four-class result?

BOUNDARYAssessment boundary
Included

Claims, metrics, source relationships, contrary evidence, limitations, population scope, version questions, authority comparison, and proof.

Not decided

Universal product efficacy, intent, deception, liability, damages, party correctness, or a legal element not established by an identified authoritative source.

02 · Canonical claim objects

Keep the claims separate.

Every number, task, explanation, and limitation remains its own object.

CLM-001“95% Sleep Staging Accuracy Compared to clinical sleep lab”
Trace

Advertising statement reproduced in the complaint

Role

Meaning presented to consumer

CLM-002“we reported Oura performance at 95% for two-class sleep/wake classification”
Trace

Oura Director of Health Science public statement

Role

Origin asserted for 95%

CLM-003“79% for four-class sleep-stage classification”
Trace

Same Oura-affiliated statement

Role

Oura-affiliated four-stage figure

CLM-004“53.18% accuracy”
Trace

Independent Scientific Reports study

Role

Four-stage result in tested clinical context

CLM-005Independent study used an older algorithm and five-minute scoring windows
Trace

Oura-affiliated public statement and 2025 letter

Role

Limitation asserted against CLM-004

03 · Evidence ledger

What it shows. What limits it.

Support and limitation travel together. A limitation narrows a finding; it does not erase it.

EV-001Complaint — 20 Aug 2026
What it shows

Reproduces 95% advertising language and alleges consumer deception.

What limits it

Allegation; not adjudicated.

EV-002Oura response — 23 Aug 2026
What it shows

States 95% refers to sleep/wake and four-stage performance is roughly 76–79%.

What limits it

Company-authored litigation response.

EV-003de Zambotti public statement — captured image supplied
What it shows

States 95% two-class, 79% four-class, and identifies 53.18% as independent/non-Oura.

What limits it

Original post URL absent from supplied capture.

EV-004Herberger et al. — 19 Mar 2025
What it shows

Reports 53.18% aggregated four-stage accuracy and kappa 0.31 in 45 sleep-clinic participants.

What limits it

Tested context does not establish all versions or populations.

EV-005de Zambotti et al. letter — 17 Sep 2025
What it shows

Criticizes version lag and five-minute comparison windows; calls for contextual reporting.

What limits it

Four of five authors employed by Oura; employees disclose equity.

EV-006Svensson et al. — 2024
What it shows

Reports binary accuracy 91.7–91.8%, sleep sensitivity 94.4–94.5%, and stage-specific accuracy 75.5–90.6%.

What limits it

Generally healthy Japanese adults; Gen3/OSSA 2.0.

EV-007Robbins et al. — 2024
What it shows

Reports 76.3% four-stage agreement and 95% sleep sensitivity.

What limits it

35 healthy adults, one night; funding and affiliation context must be preserved.

EV-008FTC authorities — cited guidance and actions
What it shows

Net-impression and claim-matched substantiation principles; Workado provides an analogous enforcement structure.

What limits it

Authority comparison only; no FTC finding against Oura is identified here.

04 · Claim support check

Does any identified evidence support 95% four-stage sleep-staging accuracy?

NONone of the identified sources establishes the advertised 95% result for Wake, Light, Deep, and REM classification.
SourceResultWhat it measuresSupports 95% staging?Plain-English meaning
Oura product representation95%Claims sleep stagingNoThe statement being tested. A marketing claim cannot serve as proof of itself.
Oura public clarification95%Sleep versus wakeNoA different two-class question, not which of four sleep stages occurred.
Oura foundational validation79%Four-stage classificationNoRelevant to staging, but 16 percentage points below 95%.
BWH validation study76.3%Four-stage classificationNoRelevant to staging, but 18.7 percentage points below 95%.
Independent clinical study53.18%Four-stage classification in the studied clinical populationNo41.82 percentage points below 95%; scoped to that population and protocol.
Oura-affiliated responseCritiques the independent study protocolNoThe critique may limit what 53.18% proves. It does not establish 95% four-stage accuracy.
05 · Detector findings

Report-derived detector results.

These are separate from the seven high-level assessment findings in the case header.

LIT-METRIC-IDENTITYObserved
Witnessed result

The 95% figure is expressly traced by Oura personnel to two-class sleep/wake classification; the challenged phrase names sleep staging.

LIT-NUMERIC-ATTRIBUTIONObserved
Witnessed result

A number generated for one identified task is attached in the challenged wording to a differently named task.

LIT-CLASS-COUNTObserved
Witnessed result

Two-class sleep/wake and four-class Wake/Light/Deep/REM are distinct output structures.

LIT-DENOMINATORNot resolved
Witnessed result

The challenged phrase does not state class prevalence, epoch denominator, participant distribution, or person-level error.

LIT-CONTRARY-EVIDENCEObserved
Witnessed result

The independent study reports 53.18% four-stage accuracy and kappa 0.31 in its tested clinical context.

LIT-LIMITATION-PRESERVATIONObserved
Witnessed result

Oura-affiliated sources identify older algorithm/version lag and five-minute windows as limitations of the 53.18% study.

LIT-VERSION-BRIDGENot established
Witnessed result

The present record does not establish a validated equivalence bridge from every cited Gen3/algorithm study to every product covered by the challenged marketing.

LIT-POPULATION-SCOPEObserved
Witnessed result

Healthy-adult validation samples and sleep-clinic populations are materially different tested populations.

LIT-FINANCIAL-INTERESTObserved
Witnessed result

The cited 2025 methodological letter discloses that four authors were Oura employees and held employee equity.

LIT-POST-CLAIM-EXPLANATIONObserved
Witnessed result

Later statements explain the 95% figure as binary sleep/wake while the challenged wording names sleep staging.

LIT-INTENTNot established
Witnessed result

The identified evidence does not establish why the challenged wording was selected.

LIT-LIABILITYNot established
Witnessed result

No court adjudication or agency finding on the challenged Oura claim is established in this record.

06 · Evidence-supported synthesis

What the record supports.

SupportedThe challenged public language presented a 95% answer to a sleep-staging accuracy question. Oura-affiliated statements trace that 95% figure to two-class sleep/wake classification and identify 79% for four-class sleep-stage classification. An independent clinical-context study reports 53.18% aggregated four-stage accuracy and kappa 0.31. Oura-affiliated sources identify version lag and five-minute comparison windows as limitations of that independent study.

What this does not establish.

  • That 53.18% is the accuracy of every Oura product, algorithm version, user, or environment.
  • That 79% or 95% applies outside the specific tasks, samples, versions, and methods from which each number arose.
  • That the challenged language was intentionally deceptive.
  • That the plaintiffs are correct on every factual or legal allegation.
  • That Oura is correct merely because it disputes the independent protocol.
  • That any court or the FTC has adjudicated this specific Oura claim.

Findings.

Party positions

Oura maintains its 95% market claim.

The independent study maintains the scope and findings of its study.

Claimants preserve pursuit of their claims.

Evidence-bounded language

The public-facing question was how accurately Oura identifies sleep stages compared with a clinical sleep lab. The advertised answer was 95%. The presently identified evidence traces the 95% figure to two-class sleep/wake classification rather than four-class sleep-stage classification. Oura-affiliated evidence identifies 79% for four-class sleep-stage classification. An independent study reports 53.18% four-class accuracy in the clinical population and protocol it tested. Oura-affiliated sources identify version lag and five-minute scoring windows as limitations of that independent result. Those limitations restrict the scope of the 53.18% finding; they do not, by themselves, establish 95% four-class sleep-stage accuracy. Sentinel does not determine liability, intent, universal product accuracy, or which litigating party is correct.

07 · Authority comparison

Compare the evidence pattern against the legal standard.

This is an authority comparison, not a government finding against Oura. The identified claim-and-evidence pattern is mapped against FTC Act § 5 substantiation principles and the documented Workado enforcement pattern. The comparison shows where the factual elements align and where they do not.
IssueOura recordFTC Workado recordComparison questionMatch
Quantified accuracy claim“95% Sleep Staging Accuracy compared to clinical sleep lab.”More than 98% accurate across all types of text.Objective performance representation made to consumers.YES
Claimed taskThe complaint says consumers were promised accurate identification of Wake, Light, Deep, and REM stages.The FTC says consumers were promised detection across the general-purpose text they actually submitted.What task did the advertisement communicate?YES
Support identified95% measures sleep versus wake; identified four-stage results are approximately 76–79%.The model was trained or fine-tuned for academic content, not the marketing content submitted by users.Does the support match the advertised task?YES
Independent result53.18% four-stage accuracy under the study protocol and population.Approximately 53% accuracy on non-academic AI-generated text.Does contrary testing materially undercut the broad percentage?YES
Exact claim substantiatedPresently identified evidence establishing 95% four-stage sleep-staging accuracy: NO.FTC alleged competent and reliable evidence establishing 98% for the advertised use: NO.Was the represented percentage supported for the represented task?NO
Fact patternHigh accuracy number attached to a broader consumer task than the supporting evidence measures.High accuracy number attached to a broader consumer task than the supporting evidence measures.Same accuracy-substantiation pattern?YES
08 · Proof & provenance

Every conclusion returns to a source object.

Captured Oura-affiliated statement · SHA-256
9944245aa7420f1e75144c711eadb7ca0ba3dca9d66cc3b0ab53e67b7ecf92d4
Canonical challenged claim · SHA-256
f5abdfb87c03fdda8d203ed68b0497efde8da258c3e8bc41884be249d72ab2eb

These hashes establish integrity only for the captured image bytes and the exact canonical claim string used in this example. They do not authenticate the original social-media account, prove authorship, or replace preservation of the original post and platform metadata.

Oura — Standing Behind Our Science (23 Aug 2026)
Surber v. Oura complaint (20 Aug 2026)
Herberger et al. — Scientific Reports (2025)
de Zambotti et al. — Sleep Advances (2025)
Robbins et al. — Sensors (2024)
Svensson et al. — Sleep Medicine (2024)
FTC Advertising Substantiation Policy
FTC Workado accuracy action (2025)
Research artifact · evaluation date 24 August 2026. The assessment reports the identified evidence record and its traceable limitations. It does not substitute for a court judgment, agency determination, or preservation of original-source metadata.