MétaCan
Menu
Back to cohort
Record W2883411829 · doi:10.2196/10527

Accuracy of Fitbit Devices: Systematic Review and Narrative Syntheses of Quantitative Data

2018· review· en· W2883411829 on OpenAlexafffundvenue
Lynne M. Feehan, Jasmina Geldman, Eric C. Sayre, Chance Park, Allison M. Ezzat, Ju Young Yoo, Clayon B. Hamilton, Linda Li

Bibliographic record

VenueJMIR mhealth and uhealth · 2018
Typereview
Languageen
FieldMedicine
TopicPhysical Activity and Health
Canadian institutionsBC Children's HospitalResearch CanadaUniversity of British Columbia
FundersCanadian Institutes of Health ResearchSimon Fraser University
KeywordsActivity trackerCINAHLApplied psychologySocial desirability biasHealth careMEDLINEActigraphyPhysical activityPsychologyComputer scienceMedicinePhysical therapyPsychological interventionSocial psychologyNursing

Abstract

fetched live from OpenAlex

BACKGROUND: Although designed as a consumer product to help motivate individuals to be physically active, Fitbit activity trackers are becoming increasingly popular as measurement tools in physical activity and health promotion research and are also commonly used to inform health care decisions. OBJECTIVE: The objective of this review was to systematically evaluate and report measurement accuracy for Fitbit activity trackers in controlled and free-living settings. METHODS: We conducted electronic searches using PubMed, EMBASE, CINAHL, and SPORTDiscus databases with a supplementary Google Scholar search. We considered original research published in English comparing Fitbit versus a reference- or research-standard criterion in healthy adults and those living with any health condition or disability. We assessed risk of bias using a modification of the Consensus-Based Standards for the Selection of Health Status Measurement Instruments. We explored measurement accuracy for steps, energy expenditure, sleep, time in activity, and distance using group percentage differences as the common rubric for error comparisons. We conducted descriptive analyses for frequency of accuracy comparisons within a ±3% error in controlled and ±10% error in free-living settings and assessed for potential bias of over- or underestimation. We secondarily explored how variations in body placement, ambulation speed, or type of activity influenced accuracy. RESULTS: We included 67 studies. Consistent evidence indicated that Fitbit devices were likely to meet acceptable accuracy for step count approximately half the time, with a tendency to underestimate steps in controlled testing and overestimate steps in free-living settings. Findings also suggested a greater tendency to provide accurate measures for steps during normal or self-paced walking with torso placement, during jogging with wrist placement, and during slow or very slow walking with ankle placement in adults with no mobility limitations. Consistent evidence indicated that Fitbit devices were unlikely to provide accurate measures for energy expenditure in any testing condition. Evidence from a few studies also suggested that, compared with research-grade accelerometers, Fitbit devices may provide similar measures for time in bed and time sleeping, while likely markedly overestimating time spent in higher-intensity activities and underestimating distance during faster-paced ambulation. However, further accuracy studies are warranted. Our point estimations for mean or median percentage error gave equal weighting to all accuracy comparisons, possibly misrepresenting the true point estimate for measurement bias for some of the testing conditions we examined. CONCLUSIONS: Other than for measures of steps in adults with no limitations in mobility, discretion should be used when considering the use of Fitbit devices as an outcome measurement tool in research or to inform health care decisions, as there are seemingly a limited number of situations where the device is likely to provide accurate measurement.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.078
metaresearch head score (Gemma)0.404
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Systematic review · Consensus signal: Systematic review
GenreCandidate signal: Review · Consensus signal: Review
Teacher disagreement score0.922
Threshold uncertainty score0.413

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0780.404
Meta-epidemiology (narrow)0.0020.002
Meta-epidemiology (broad)0.0090.008
Bibliometrics0.0270.027
Science and technology studies0.0010.003
Scholarly communication0.0060.008
Open science0.0030.004
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0070.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.385
GPT teacher head0.535
Teacher spread0.150 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSystematic review
DomainMethods
GenreReview

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations665
Published2018
Admission routes3
Has abstractyes

Explore more

Same venueJMIR mhealth and uhealthSame topicPhysical Activity and HealthFrench-language works237,207