Accuracy and Precision of Consumer-Grade Wearable Activity Monitors for Assessing Time Spent in Sedentary Behavior in Children and Adolescents: Systematic Review
Bibliographic record
Abstract
BACKGROUND: A large number of wearable activity monitor models are released and used each year by consumers and researchers. As more studies are being carried out on children and adolescents in terms of sedentary behavior (SB) assessment, knowledge about accurate and precise monitoring devices becomes increasingly important. OBJECTIVE: The main aim of this systematic review was to investigate and communicate findings on the accuracy and precision of consumer-grade physical activity monitors in assessing the time spent in SB in children and adolescents. METHODS: Searches of PubMed (MEDLINE), Scopus, SPORTDiscus (full text), ProQuest, Open Access Theses and Dissertations, DART Europe E-theses Portal, and Networked Digital Library of Theses and Dissertations electronic databases were performed. All relevant studies that compared different types of consumer-grade monitors using a comparison method in the assessment of SB, published in European languages from 2015 onward were considered for inclusion. The risk of bias was estimated using Consensus-Based Standards for the Selection of Health Status Measurement Instruments. For enabling comparisons of accuracy measures within the studied outcome domain, measurement accuracy interpretation was based on group mean or percentage error values and 90% CI. Acceptable limits were predefined as -10% to +10% error in controlled and free-living settings. For determining the number of studies with group error percentages that fall within or outside one of the sides from previously defined acceptable limits, two 1-sided tests of equivalence were carried out, and the direction of measurement error was examined. RESULTS: , which represents the percentage of total variation across studies due to heterogeneity, amounted to 94%. The summary effect size based on the random effects model was not statistically significant (effect size=14.36, SE 12.04, 90% CI -5.45 to 34.17; P=.23). According to the equivalence test results, consumer-grade physical activity monitors did not generate equivalent estimates of SB in relation to the comparison methods. Majority of the studies (3/7, 43%) that reported the mean absolute percentage errors have reported values of <30%. CONCLUSIONS: This is the first study that has attempted to synthesize available evidence on the accuracy and precision of consumer-grade physical activity monitors in measuring SB in children and adolescents. We found very few studies on the accuracy and almost no evidence on the precision of wearable activity monitors. The presented results highlight the large heterogeneity in this area of research. TRIAL REGISTRATION: PROSPERO CRD42021251922; https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=251922.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.040 | 0.206 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.012 | 0.013 |
| Bibliometrics | 0.011 | 0.011 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".