Automatic Identification of Physical Activity Type and Duration by Wearable Activity Trackers: A Validation Study
Bibliographic record
Abstract
BACKGROUND: Activity trackers are now ubiquitous in certain populations, with potential applications for health promotion and monitoring and chronic disease management. Understanding the accuracy of this technology is critical to the appropriate and productive use of wearables in health research. Although other peer-reviewed validations have examined other features (eg, steps and heart rate), no published studies to date have addressed the accuracy of automatic activity type detection and duration accuracy in wearable trackers. OBJECTIVE: The aim of this study was to examine the ability of 4 commercially available wearable activity trackers (Fitbits Flex 2, Fitbit Alta HR, Fitbit Charge 2, and Garmin Vívosmart HR), in a controlled setting, to correctly and automatically identify the type and duration of the physical activity being performed. METHODS: A total of 8 activity types, including walking and running (on both a treadmill and outdoors), a run embedded in walking bouts, elliptical use, outdoor biking, and pool lap swimming, were tested by 28 to 34 healthy adult participants (69 total participants who participated in some to all activity types). Actual activity type and duration were recorded by study personnel and compared with tracker data using descriptive statistics and mean absolute percent error (MAPE). RESULTS: The proportion of trials in which the activity type was correctly identified was 93% to 97% (depending on the tracker) for treadmill walking, 93% to 100% for treadmill running, 36% to 62% for treadmill running when preceded and followed by a walk, 97% to 100% for outdoor walking, 100% for outdoor running, 3% to 97% for using an elliptical, 44% to 97% for biking, and 87.5% for swimming. When activities were correctly identified, the MAPE of the detected duration versus the actual activity duration was between 7% and 7.9% for treadmill walking, 8.7% and 144.8% for treadmill running, 23.6% and 28.9% for treadmill running when preceded and followed by a walk, 4.9% and 11.8% for outdoor walking, 5.6% and 9.6% for outdoor running, 9.7% and 13% for using an elliptical, 9.5% and 17.7% for biking, and was 26.9% for swimming. CONCLUSIONS: In a controlled setting, wearable activity trackers provide accurate recognition of the type of some common physical activities, especially outdoor walking and running and walking on a treadmill. The accuracy of measurement of activity duration varied considerably by activity type and tracker model and was poor for complex sets of activity, such as a run embedded within 2 walking segments.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".