State of the science and recommendations for using wearable technology in sleep and circadian research
Bibliographic record
Abstract
Wearable sleep-tracking technology is of growing use in the sleep and circadian fields, including for applications across other disciplines, inclusive of a variety of disease states. Patients increasingly present sleep data derived from their wearable devices to their providers and the ever-increasing availability of commercial devices and new-generation research/clinical tools has led to the wide adoption of wearables in research, which has become even more relevant given the discontinuation of the Philips Respironics Actiwatch. Standards for evaluating the performance of wearable sleep-tracking devices have been introduced and the available evidence suggests that consumer-grade devices exceed the performance of traditional actigraphy in assessing sleep as defined by polysomnogram. However, clear limitations exist, for example, the misclassification of wakefulness during the sleep period, problems with sleep tracking outside of the main sleep bout or nighttime period, artifacts, and unclear translation of performance to individuals with certain characteristics or comorbidities. This is of particular relevance when person-specific factors (like skin color or obesity) negatively impact sensor performance with the potential downstream impact of augmenting already existing healthcare disparities. However, wearable sleep-tracking technology holds great promise for our field, given features distinct from traditional actigraphy such as measurement of autonomic parameters, estimation of circadian features, and the potential to integrate other self-reported, objective, and passively recorded health indicators. Scientists face numerous decision points and barriers when incorporating traditional actigraphy, consumer-grade multi-sensor devices, or contemporary research/clinical-grade sleep trackers into their research. Considerations include wearable device capabilities and performance, target population and goals of the study, wearable device outputs and availability of raw and aggregate data, and data extraction, processing, and analysis. Given the difficulties in the implementation and utilization of wearable sleep-tracking technology in real-world research and clinical settings, the following State of the Science review requested by the Sleep Research Society aims to address the following questions. What data can wearable sleep-tracking devices provide? How accurate are these data? What should be taken into account when incorporating wearable sleep-tracking devices into research? These outstanding questions and surrounding considerations motivated this work, outlining practical recommendations for using wearable technology in sleep and circadian research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.060 | 0.124 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.006 |
| Bibliometrics | 0.008 | 0.009 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.011 | 0.012 |
| Open science | 0.007 | 0.007 |
| Research integrity | 0.015 | 0.018 |
| Insufficient payload (model declined to judge) | 0.031 | 0.022 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".