Issue 155  /  August 25, 2026  /  Feature

Everyone's Measuring. Nobody's Built the Next Step

The largest evidence review ever run on wearables found the device alone performs about 50% worse than the device inside a programme. Femtech is funding the device. The programme is the part that needs a clinician, a diagnosis, and a billable code.

Everyone's Measuring. Nobody's Built the Next Step

The best available answer to whether tracking improves health comes from an umbrella review. Ferguson and colleagues at the University of South Australia searched seven databases through April 2021 and pooled 39 systematic reviews and meta-analyses covering 163,992 participants across clinical and non-clinical populations, published in The Lancet Digital Health in August 2022.

Wearable activity trackers improved physical activity at standardised mean differences of 0.3 to 0.6, body composition at 0.7 to 2.0 and fitness at 0.3. In plain terms, roughly 1,800 additional steps per day, about 40 minutes more walking, and around 1 kg of weight reduction. The authors call the benefit clinically important and durable to at least six months.

Then the second half. Effects on blood pressure, cholesterol and glycosylated haemoglobin were small and often non-significant. Cholesterol came in at SMD −0.06, 95% CI −0.31 to 0.19. Diastolic blood pressure at −0.1, −0.28 to 0.10. Quality of life showed little evidence of effect across four meta-analyses.

The authors are careful here, and it matters. They note that effect sizes of 0.05 to 0.2 are common in medical research and can still be meaningful, that most component trials ran three months or less, and that some readers may find their own reading of the physiological outcomes overly conservative.

Thirty-four of the 39 reviews were rated critically low confidence on AMSTAR 2, though sensitivity analysis found results consistent when the stronger reviews were isolated.

Buried in the discussion, Ferguson's team flags that few of the 39 reviews tested trackers on their own. Brickwood and colleagues did in 2019, separately meta-analysing 16 multifaceted interventions against seven that used the tracker alone. The multifaceted interventions produced effects around 50% larger.

There is sufficient evidence to recommend wearable activity trackers at least as an adjunct to programmes aiming to increase physical activity.

An adjunct. Not the intervention.

That is the whole argument for what happens next in women's health hardware. Measurement is the fundable layer. It ships, it demos, it produces a number every morning and a subscription every month.

Conversion, meaning the step that turns a reading into treatment somebody pays for, is not fundable in the same way, because it requires a clinician, a diagnosis and a billable code.

Femtech has been building the first layer at speed on the assumption that the second follows. The strongest evidence in the category says the first layer is roughly half as effective without the second, in the one domain where measurement is easiest.

Because step counting is the most favourable case the sector has. A pedometer counts. There is no inference layer, no algorithm mapping a peripheral proxy onto a hidden physiological state, no dispute about what the number represents. And the action implied by the number is obvious, available and free. Walk more.

Every biomarker femtech is now selling sits on the other side of both conditions. The sensor infers rather than counts, and the implied action is a prescription rather than a walk.

Which brings us to the part of this that ends up in court.

Madison Surber filed a proposed class action against Oura Inc. and Oura Health Oy on August 20 in the Northern District of California, docketed 3:26-cv-08686, represented by Clarkson Law Firm.

She purchased an Oura Ring 4 Gold for approximately $513.68 in May 2025. The complaint challenges the phrases "Built for accuracy" and "Unparalleled Accuracy" alongside advertised figures of 79% and 95%, arguing a finger-worn device without EEG, EOG or EMG sensors cannot support them.

Seven causes of action. Oura disputes the allegations and has said it will defend them.

This is the third time this case has been filed in that courthouse.

In May 2015, James Brickman sued Fitbit over the sleep-tracking function on the Flex, One and Ultra, alleging the accelerometer could not deliver the advertised sleep quality data, and that buyers paid roughly $30 extra for it. Brickman v. Fitbit, 3:15-cv-02077-JD. Same statutes: California's Unfair Competition Law, False Advertising Law and Consumers Legal Remedies Act, plus warranty and misrepresentation claims.

In January 2016, McLellan v. Fitbit, 3:16-cv-00036-JD, made the parallel argument about PurePulse heart rate.

Fitbit moved to dismiss Brickman by attacking the plaintiffs' science and submitting a compilation of studies it said validated accelerometer-based sleep tracking. Judge James Donato declined to consider them, on the grounds that a motion to dismiss tests the sufficiency of the complaint rather than the merits.

He denied dismissal in July 2016, certified a class in November 2017, and the case ultimately settled, with roughly $7M in class counsel fees approved.

Fitbit's public position throughout was that its trackers were not intended to be scientific or medical devices.

Eleven years later, the same district, the same statutes, the same defence.

The technical dispute in the Oura case is narrower than the coverage suggests, and the company's response is worth reading because it clarifies something the category has been vague about for a decade.

In a post published August 23 and expert-reviewed by Shyamal Patel, PhD, its SVP of Science, Oura explains how epoch-level accuracy is computed and states that sleep versus wake detection reaches 90% to 96% agreement with polysomnography, and that this is what the 95% figure refers to.

Four-stage classification, distinguishing light, deep, REM and wake, is described in the same post as reaching roughly 76% to 79% in healthy adults.

Robbins and colleagues at Brigham and Women's Hospital reported 76.3% four-stage and 92% two-stage agreement in 2024, rating Oura the most accurate consumer tracker tested. Ghorbani and colleagues at the National University of Singapore found 76.4% across 157 nights in 2022.

Svensson and colleagues at the University of Tokyo evaluated 96 participants across 421,045 epochs and reported two-stage accuracy of 91.7% to 91.8%, with specificity of 73.0% to 74.6%. The foundational 2021 Sensors study reporting 79% was authored by Altini and Kinnunen, both Oura-affiliated, which the company discloses.

Oura also makes a point the category should absorb. Agreement between two trained humans scoring the same polysomnography recording runs around 83%, and falls further in people with sleep disorders.

The reference standard is itself an estimate, so no device can reach 100%, and the 53.18% four-stage figure now circulating comes from exactly the population where the reference degrades.

That result is Herberger and colleagues in Scientific Reports, March 2025, from the sleep laboratory at Charité University Medicine Berlin, in patients carrying diagnoses including obstructive and central apnoea, chronic insomnia, restless legs and narcolepsy, mean age 54.6, mean BMI 33.5.

Oura returned usable data on 31 of 45 nights. Single night, first wear, no calibration period, three rings under a full montage. The authors state all of it.

So the ring measures roughly what a finger sensor can be expected to measure. The dispute is about which of two numbers gets printed next to a four-stage word in an advertisement, and that is a question about claim architecture rather than engineering.

Claim architecture is where the regulatory structure sits, and it is the part with direct consequences for women's health.

Consumer wellness devices are validated signal by signal and marketed device by device. Nothing in the general wellness pathway requires those two to align.

The same Oura sensor stack demonstrates this cleanly. Its temperature output is regulated, cleared as K202897 on June 24, 2021 under 21 CFR 884.5370, Software Application for Contraception, Class II, product code PYT.

The clearance added Oura as an alternative source of daily basal body temperature for the Natural Cycles algorithm, which was not itself modified.

The supporting study enrolled 40 women who already used Natural Cycles with an oral thermometer, mean age 31.3, 38 of them in Sweden, across 223 cycles with at least one positive LH test in 87 complete cycles. The finding was that the algorithm identified ovulation from either input, and that Oura's input produced 1.6 additional non-fertile days in the luteal phase without raising pregnancy risk.

Forty women, one input substitution, one indication. Sleep staging sits outside it. So does perimenopause.

On Oura's own Natural Cycles page, last modified April 28, 2026, the temperature sensor's lab accuracy and the sleep staging algorithm's performance against polysomnography appear in a single sentence. Both claims are permitted. One was reviewed by a regulator and one was not, and the page is under no obligation to distinguish them.

Which returns to the harder question, the one the accuracy debate never reaches. A woman has the number. What is she supposed to do with it?

Oura addresses a version of this directly, responding to members who feel unrested despite a good score by arguing that subjective experience and objective measurement are separate dimensions and a mismatch is not evidence of inaccuracy. That is defensible.

It also leaves out what the number does on its own, which has been tested.

Gavriloff and colleagues published the experiment in the Journal of Sleep Research in 2018. Sixty-three adults meeting DSM-5 criteria for insomnia disorder were randomised to receive fabricated sleep-efficiency feedback, positive or negative, delivered at rise-time through an actigraphy watch built to simulate a consumer wearable. The negative-feedback group reported decreased alert cognition by evening at d = 0.79 and increased sleepiness and fatigue at d = 0.55. On objective psychomotor vigilance the groups did not differ, d = 0.12, nor on sleep-related attentional bias, d = 0.20. Participants were people with diagnosed insomnia rather than general consumers, the feedback was single-instance, and participants were debriefed and offered access to Sleepio, a commercial digital CBT programme with which two authors are affiliated.

Within those limits the result is specific. A number corresponding to nothing changed how people said they felt, substantially, and did not change how they performed. The output acts as an intervention whether or not it is accurate, and it acts on perception rather than physiology.

Some women have arrived at that conclusion without the study.

De Boer and colleagues at Tilburg University interviewed 13 menopausal women in the Netherlands for a paper published in Health in 2024. Most had stopped, sharply reduced or actively resisted self-tracking during menopause, describing bodies they experienced as knowledgeable rather than as objects requiring measurement.

Thirteen women, one country, qualitative, none of them on hormone therapy. It proves nothing on its own. It is also one of very few studies that has asked the question at all, and the thinness of that literature is itself worth noticing while capital moves the other way.

Clair Health raised an $11.6M seed led by Khosla Ventures in April, with a16z speedrun and Anne Wojcicki participating, for a wrist-worn device the company says infers estrogen, progesterone, LH and FSH continuously from ten biosensors including what it describes as a novel biomagnetic sensor.

Founding members pay $369 plus $9.99 monthly, with shipping stated for December 2026. Clair describes the product as a wellness device rather than a diagnostic. Mira, Inne and Eli Health hold adjacent positions on varying regulatory footing.

Set the evidence base beside the one the sleep companies are working from. Oura says its staging algorithms were trained on more than 1,200 nights of clinical polysomnography. Four independent academic groups plus one in-house team have published validation against a defined reference standard, using a scoring manual that has existed for decades and an inter-rater benchmark that is itself measurable.

Continuous non-invasive hormone inference has none of that. No scoring manual, no inter-rater benchmark, no cleared predicate, no external reference against which a claim could be checked.

A company can be entirely truthful about a signal that nobody is yet in a position to verify.

And the actionability problem gets worse rather than better as the biomarker gets more specific. A step count implies walking. An estradiol curve implies a prescription, which requires a clinician, a diagnosis and a payable code.

That is the same wall the menopause tracking market hit. The measurement layer is fundable, buildable and shippable. The layer that converts a reading into care is none of those things.

The Luteal read, for anyone building or investing in this category: Ask which specific signal was validated, against which reference standard, in how many people and of what age. "FDA-cleared" attaches to a signal and an indication, not to a device.

Forty women in Sweden cleared a temperature input for contraception, and that clearance says nothing about anything else the same ring displays.

Ask whether the marketing sentence describes the validated signal or a neighboring one. That gap is what eleven years of Northern District of California filings have been about, and Brickman shows that a compilation of supporting studies does not resolve it at the pleading stage.

Ask what the user does differently on Tuesday because of the number, and who pays for it.

Ferguson's 163,992 participants are the benchmark.

In the easiest domain in the category, the device on its own was about half as effective as the device inside a programme, and the authors recommended it as an adjunct.

The companies that win the next cycle will not be the ones with the better sensor.

They will be the ones that own the thing the number is supposed to trigger.

The Luteal covers the business, science and policy of women's health. Nothing here is medical advice.