Validity of Apple Watch, Garmin Forerunner<sup>®</sup> 935 and GENEActiv for estimating energy expenditure during close quarter battle training in Special Forces soldiers
Bibliographic record
Abstract
Abstract The purpose of this study was to investigate the validity of three wrist‐worn devices for estimating energy expenditure (EE) and heart rate (HR) during close “Close Quarter Battle” (CQB). Fifty male soldiers (mean ± SD: age 30.9 ± 4.6 years, height: 1.81 ± 0.64 m and body mass 87.3 ± 7.7 kg) wore three activity monitors (Apple Watch 5, Garmin Forerunner ® 935 and GENEActiv accelerometer), a Metamax 3B metabolic cart and a Polar chest strap, whilst conducting a CQB training activity (duration: 26.6 ± 5.0 min). EE and HR data from each test device were compared against criterion measures using ordinary least products regression, 95% limits of agreement, equivalence testing and Mean Absolute Percentage Error (MAPE). Based upon the criterion measure the mean EE for the activity was 372.2 ± 57.6 kcal. All of the devices tested demonstrated fixed and/or proportional bias for EE and a MAPE of >10% (Apple 11.3%, Garmin 15.3%, GENEActiv 57.7%) and therefore did not agree with the criterion. The Apple Watch was a valid method for measuring HR, with a MAPE of 0.6%, and differences with the criterion falling within acceptable limits (≤1 bpm = 83.7%; ≤3 bpm = 97.5% and ≤5 bpm = 97.5%), whereas the Garmin Forerunner ® 935 was not valid for measuring HR due to an unacceptable difference compared to the criterion (≤1 bpm = 19.1%; ≤3 bpm = 33.3% and ≤5 bpm = 45.2%). Overall, the Apple Watch 5 can be recommended for measuring HR, but none of the devices are recommended for estimating EE, during CQB.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".