Measuring Free-Living Physical Activity With Three Commercially Available Activity Monitors for Telemonitoring Purposes: Validation Study
Bibliographic record
Abstract
BACKGROUND: Remote monitoring of physical activity in patients with chronic conditions could be useful to offer care professionals real-time assessment of their patient's daily activity pattern to adjust appropriate treatment. However, the validity of commercially available activity trackers that can be used for telemonitoring purposes is limited. OBJECTIVE: The purpose of this study was to test usability and determine the validity of 3 consumer-level activity trackers as a measure of free-living activity. METHODS: A usability evaluation (study 1) and validation study (study 2) were conducted. In study 1, 10 individuals wore one activity tracker for a period of 30 days and filled in a questionnaire on ease of use and wearability. In study 2, we validated three selected activity trackers (Apple Watch, Misfit Shine, and iHealth Edge) and a fourth pedometer (Yamax Digiwalker) against the reference standard (Actigraph GT3X) in 30 healthy participants for 72 hours. Outcome measures were 95% limits of agreement (LoA) and bias (Bland-Altman analysis). Furthermore, median absolute differences (MAD) were calculated. Correction for bias was estimated and validated using leave-one-out cross validation. RESULTS: Usability evaluation of study 1 showed that iHealth Edge and Apple Watch were more comfortable to wear as compared with the Misfit Flash. Therefore, the Misfit Flash was replaced by Misfit Shine in study 2. During study 2, the total number of steps of the reference standard was 21,527 (interquartile range, IQR 17,475-24,809). Bias and LoA for number of steps from the Apple Watch and iHealth Edge were 968 (IQR -5478 to 7414) and 2021 (IQR -4994 to 9036) steps. For Misfit Shine and Yamax Digiwalker, bias was -1874 and 2004, both with wide LoA of (13,869 to 10,121) and (-10,932 to 14,940) steps, respectively. The Apple Watch noted the smallest MAD of 7.7% with the Actigraph, whereas the Yamax Digiwalker noted the highest MAD (20.3%). After leave-one-out cross validation, accuracy estimates of MAD of the iHealth Edge and Misfit Shine were within acceptable limits with 10.7% and 11.3%, respectively. CONCLUSIONS: Overall, the Apple Watch and iHealth Edge were positively evaluated after wearing. Validity varied widely between devices, with the Apple Watch being the most accurate and Yamax Digiwalker the least accurate for step count in free-living conditions. The iHealth Edge underestimates number of steps but can be considered reliable for activity monitoring after correction for bias. Misfit Shine overestimated number of steps and cannot be considered suitable for step count because of the low agreement. Future studies should focus on the added value of remotely monitoring activity patterns over time in chronic patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".