Validation of Sleep Measurements of an Actigraphy Watch: Instrument Validation Study
Bibliographic record
Abstract
BACKGROUND: The iAide2 (Tokai) physical activity monitoring system includes diverse measurements and wireless features useful to researchers. The iAide2's sleep measurement capabilities have not been compared to validated sleep measurement standards in any published work. OBJECTIVE: We aimed to assess the iAide2's sleep duration and total sleep time (TST) measurement performance and perform calibration if needed. METHODS: We performed free-living sleep monitoring in 6 convenience-sampled participants without known sleep disorders recruited from within the Waki DTx Laboratory at the Graduate School of Medicine, University of Tokyo. To assess free-living sleep, we validated the iAide2 against a second actigraph that was previously validated against polysomnography, the MotionWatch 8 (MW8; CamNtech Ltd). The participants wore both devices on the nondominant arm, with the MW8 closest to the hand, all day except when bathing. The MW8 and iAide2 assessments both used the MW8 EVENT-marker button to record bedtime and risetime. For the MW8, MotionWare Software (version 1.4.20; CamNtech Ltd) provided TST, and we calculated sleep duration from the sleep onset and sleep offset provided by the software. We used a similar process with the iAide2, using iAide2 software (version 7.0). We analyzed 64 nights and evaluated the agreement between the iAide2 and the MW8 for sleep duration and TST based on intraclass correlation coefficients (ICCs). RESULTS: The absolute ICCs (2-way mixed effects, absolute agreement, single measurement) for sleep duration (0.69, 95% CI -0.07 to 0.91) and TST (0.56, 95% CI -0.07 to 0.82) were moderate. The consistency ICC (2-way mixed effects, consistency, single measurement) was excellent for sleep duration (0.91, 95% CI 0.86-0.95) and moderate for TST (0.78, 95% CI 0.67-0.86). We determined a simple calibration approach. After calibration, the ICCs improved to 0.96 (95% CI 0.94-0.98) for sleep duration and 0.82 (95% CI 0.71-0.88) for TST. The results were not sensitive to the specific participants included, with an ICC range of 0.96-0.97 for sleep duration and 0.79-0.87 for TST when applying our calibration equation to data removing one participant at a time and 0.96-0.97 for sleep duration and 0.79-0.86 for TST when recalibrating while removing one participant at a time. CONCLUSIONS: The measurement errors of the uncalibrated iAide2 for both sleep duration and TST seem too large for them to be useful as absolute measurements, though they could be useful as relative measurements. The measurement errors after calibration are low, and the calibration approach is general and robust, validating the use of iAide2's sleep measurement functions alongside its other features in physical activity research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.038 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".