Physical Activity Assessment Between Consumer- and Research-Grade Accelerometers: A Comparative Study in Free-Living Conditions
Bibliographic record
Abstract
BACKGROUND: Wearable activity monitors such as Fitbit enable users to track various attributes of their physical activity (PA) over time and have the potential to be used in research to promote and measure PA behavior. However, the measurement accuracy of Fitbit in absolute free-living conditions is largely unknown. OBJECTIVE: To examine the measurement congruence between Fitbit Flex and ActiGraph GT3X for quantifying steps, metabolic equivalent tasks (METs), and proportion of time in sedentary activity and light-, moderate-, and vigorous-intensity PA in healthy adults in free-living conditions. METHODS: A convenience sample of 19 participants (4 men and 15 women), aged 18-37 years, concurrently wore the Fitbit Flex (wrist) and ActiGraph GT3X (waist) for 1- or 2-week observation periods (n=3 and n=16, respectively) that included self-reported bouts of daily exercise. Data were examined for daily activity, averaged over 14 days and for minutes of reported exercise. Average day-level data included steps, METs, and proportion of time in different intensity levels. Minute-level data included steps, METs, and mean intensity score (0 = sedentary, 3 = vigorous) for overall reported exercise bouts (N=120) and by exercise type (walking, n=16; run or sports, n=44; cardio machine, n=20). RESULTS: Measures of steps were similar between devices for average day- and minute-level observations (all P values > .05). Fitbit significantly overestimated METs for average daily activity, for overall minutes of reported exercise bouts, and for walking and run or sports exercises (mean difference 0.70, 1.80, 3.16, and 2.00 METs, respectively; all P values < .001). For average daily activity, Fitbit significantly underestimated the proportion of time in sedentary and light intensity by 20% and 34%, respectively, and overestimated time by 3% in both moderate and vigorous intensity (all P values < .001). Mean intensity scores were not different for overall minutes of exercise or for run or sports and cardio-machine exercises (all P values > .05). CONCLUSIONS: Fitbit Flex provides accurate measures of steps for daily activity and minutes of reported exercise, regardless of exercise type. Although the proportion of time in different intensity levels varied between devices, examining the mean intensity score for minute-level bouts across different exercise types enabled interdevice comparisons that revealed similar measures of exercise intensity. Fitbit Flex is shown to have measurement limitations that may affect its potential utility and validity for measuring PA attributes in free-living conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".