Measurement of Heart Rate Using the Withings ScanWatch Device During Free-living Activities: Validation Study
Bibliographic record
Abstract
BACKGROUND: Wrist-worn devices that incorporate photoplethysmography (PPG) sensing represent an exciting means of measuring heart rate (HR). A number of studies have evaluated the accuracy of HR measurements produced by these devices in controlled laboratory environments. However, it is also important to establish the accuracy of measurements produced by these devices outside the laboratory, in real-world, consumer use conditions. OBJECTIVE: This study sought to examine the accuracy of HR measurements produced by the Withings ScanWatch during free-living activities. METHODS: A sample of convenience of 7 participants volunteered (3 male and 4 female; mean age 64, SD 10 years; mean height 164, SD 4 cm; mean weight 77, SD 16 kg) to take part in this real-world validation study. Participants were instructed to wear the ScanWatch for a 12-hour period on their nondominant wrist as they went about their day-to-day activities. A Polar H10 heart rate sensor was used as the criterion measure of HR. Participants used a study diary to document activities undertaken during the 12-hour study period. These activities were classified according to the 11 following domains: desk work, eat or drink, exercise, gardening, household activities, self-care, shopping, sitting, sleep, travel, and walking. Validity was assessed using the Bland-Altman analysis, concordance correlation coefficient (CCC), and mean absolute percentage error (MAPE). RESULTS: Across all activity domains, the ScanWatch measured HR with MAPE values <10%, except for the shopping activity domain (MAPE=10.8%). The activity domains that were more sedentary in nature (eg, desk work, eat or drink, and sitting) produced the most accurate HR measurements with a small mean bias and MAPE values <5%. Moderate to strong correlations (CCC=0.526-0.783) were observed between devices for all activity domains, except during the walking activity domain, which demonstrated a weak correlation (CCC=0.164) between devices. CONCLUSIONS: The results of this study show that the ScanWatch measures HR with a degree of accuracy that is acceptable for general consumer use; however, it would not be suitable in circumstances where more accurate measurements of HR are required, such as in health care or in clinical trials. Overall, the ScanWatch was less accurate at measuring HR during ambulatory activities (eg, walking, gardening, and household activities) compared to more sedentary activities (eg, desk work, eat or drink, and sitting). Further larger-scale studies examining this device in different populations and during different activities are required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".