Physical Activity Assessment Between Consumer- and Research-Grade Accelerometers: A Comparative Study in Free-Living Conditions
Bibliographic record
Abstract
BACKGROUND: Wearable activity monitors such as Fitbit enable users to track various attributes of their physical activity (PA) over time and have the potential to be used in research to promote and measure PA behavior. However, the measurement accuracy of Fitbit in absolute free-living conditions is largely unknown. OBJECTIVE: To examine the measurement congruence between Fitbit Flex and ActiGraph GT3X for quantifying steps, metabolic equivalent tasks (METs), and proportion of time in sedentary activity and light-, moderate-, and vigorous-intensity PA in healthy adults in free-living conditions. METHODS: A convenience sample of 19 participants (4 men and 15 women), aged 18-37 years, concurrently wore the Fitbit Flex (wrist) and ActiGraph GT3X (waist) for 1- or 2-week observation periods (n=3 and n=16, respectively) that included self-reported bouts of daily exercise. Data were examined for daily activity, averaged over 14 days and for minutes of reported exercise. Average day-level data included steps, METs, and proportion of time in different intensity levels. Minute-level data included steps, METs, and mean intensity score (0 = sedentary, 3 = vigorous) for overall reported exercise bouts (N=120) and by exercise type (walking, n=16; run or sports, n=44; cardio machine, n=20). RESULTS: Measures of steps were similar between devices for average day- and minute-level observations (all P values > .05). Fitbit significantly overestimated METs for average daily activity, for overall minutes of reported exercise bouts, and for walking and run or sports exercises (mean difference 0.70, 1.80, 3.16, and 2.00 METs, respectively; all P values < .001). For average daily activity, Fitbit significantly underestimated the proportion of time in sedentary and light intensity by 20% and 34%, respectively, and overestimated time by 3% in both moderate and vigorous intensity (all P values < .001). Mean intensity scores were not different for overall minutes of exercise or for run or sports and cardio-machine exercises (all P values > .05). CONCLUSIONS: Fitbit Flex provides accurate measures of steps for daily activity and minutes of reported exercise, regardless of exercise type. Although the proportion of time in different intensity levels varied between devices, examining the mean intensity score for minute-level bouts across different exercise types enabled interdevice comparisons that revealed similar measures of exercise intensity. Fitbit Flex is shown to have measurement limitations that may affect its potential utility and validity for measuring PA attributes in free-living conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".