Sleep Tracking of a Commercially Available Smart Ring and Smartwatch Against Medical-Grade Actigraphy in Everyday Settings: Instrument Validation Study
Bibliographic record
Abstract
BACKGROUND: Assessment of sleep quality is essential to address poor sleep quality and understand changes. Owing to the advances in the Internet of Things and wearable technologies, sleep monitoring under free-living conditions has become feasible and practicable. Smart rings and smartwatches can be employed to perform mid- or long-term home-based sleep monitoring. However, the validity of such wearables should be investigated in terms of sleep parameters. Sleep validation studies are mostly limited to short-term laboratory tests; there is a need for a study to assess the sleep attributes of wearables in everyday settings, where users engage in their daily routines. OBJECTIVE: This study aims to evaluate the sleep parameters of the Oura ring along with the Samsung Gear Sport watch in comparison with a medically approved actigraphy device in a midterm everyday setting, where users engage in their daily routines. METHODS: We conducted home-based sleep monitoring in which the sleep parameters of 45 healthy individuals (23 women and 22 men) were tracked for 7 days. Total sleep time (TST), sleep efficiency (SE), and wake after sleep onset (WASO) of the ring and watch were assessed using paired t tests, Bland-Altman plots, and Pearson correlation. The parameters were also investigated considering the gender of the participants as a dependent variable. RESULTS: We found significant correlations between the ring's and actigraphy's TST (r=0.86; P<.001), WASO (r=0.41; P<.001), and SE (r=0.47; P<.001). Comparing the watch with actigraphy showed a significant correlation in TST (r=0.59; P<.001). The mean differences in TST, WASO, and SE of the ring and actigraphy were within satisfactory ranges, although there were significant differences between the parameters (P<.001); TST and SE mean differences were also within satisfactory ranges for the watch, and the WASO was slightly higher than the range (31.27, SD 35.15). However, the mean differences of the parameters between the watch and actigraphy were considerably higher than those of the ring. The watch also showed a significant difference in TST (P<.001) between female and male groups. CONCLUSIONS: In a sample population of healthy adults, the sleep parameters of both the Oura ring and Samsung watch have acceptable mean differences and indicate significant correlations with actigraphy, but the ring outperforms the watch in terms of the nonstaging sleep parameters.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".