Accuracy of Apple Watch Measurements for Heart Rate and Energy Expenditure in Patients With Cardiovascular Disease: Cross-Sectional Study
Bibliographic record
Abstract
BACKGROUND: Wrist-worn tracking devices such as the Apple Watch are becoming more integrated in health care. However, validation studies of these consumer devices remain scarce. OBJECTIVES: This study aimed to assess if mobile health technology can be used for monitoring home-based exercise in future cardiac rehabilitation programs. The purpose was to determine the accuracy of the Apple Watch in measuring heart rate (HR) and estimating energy expenditure (EE) during a cardiopulmonary exercise test (CPET) in patients with cardiovascular disease. METHODS: Forty patients (mean age 61.9 [SD 15.2] yrs, 80% male) with cardiovascular disease (70% ischemic, 22.5% valvular, 7.5% other) completed a graded maximal CPET on a cycle ergometer while wearing an Apple Watch. A 12-lead electrocardiogram (ECG) was used to measure HR; indirect calorimetry was used for EE. HR was analyzed at three levels of intensity (seated rest, HR1; moderate intensity, HR2; maximal performance, HR3) for 30 seconds. The EE of the entire test was used. Bias or mean difference (MD), standard deviation of difference (SDD), limits of agreement (LoA), mean absolute error (MAE), mean absolute percentage error (MAPE), and intraclass correlation coefficients (ICCs) were calculated. Bland-Altman plots and scatterplots were constructed. RESULTS: SDD for HR1, HR2, and HR3 was 12.4, 16.2, and 12.0 bpm, respectively. Bias and LoA (lower, upper LoA) were 3.61 (-20.74, 27.96) for HR1, 0.91 (-30.82, 32.63) for HR2, and -1.82 (-25.27, 21.63) for HR3. MAE was 6.34 for HR1, 7.55 for HR2, and 6.90 for HR3. MAPE was 10.69% for HR1, 9.20% for HR2, and 6.33% for HR3. ICC was 0.729 (P<.001) for HR1, 0.828 (P<.001) for HR2, and 0.958 (P<.001) for HR3. Bland-Altman plots and scatterplots showed good correlation without systematic error when comparing Apple Watch with ECG measurements. SDD for EE was 17.5 kcal. Bias and LoA were 30.47 (-3.80, 64.74). MAE was 30.77; MAPE was 114.72%. ICC for EE was 0.797 (P<.001). The Bland-Altman plot and a scatterplot directly comparing Apple Watch and indirect calorimetry showed systematic bias with an overestimation of EE by the Apple Watch. CONCLUSIONS: In patients with cardiovascular disease, the Apple Watch measures HR with clinically acceptable accuracy during exercise. If confirmed, it might be considered safe to incorporate the Apple Watch in HR-guided training programs in the setting of cardiac rehabilitation. At this moment, however, it is too early to recommend the Apple Watch for cardiac rehabilitation. Also, the Apple Watch systematically overestimates EE in this group of patients. Caution might therefore be warranted when using the Apple Watch for measuring EE.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".