The quality of physical activity guidelines, but not the specificity of their recommendations, has improved over time: a systematic review and critical appraisal
Bibliographic record
Abstract
While numerous guidelines for the prescription of physical activity are released each year, the quality and practicability of these guidelines is unknown. We assessed the quality of 95 guidance documents published since 2000 that included recommendations about physical activity for the promotion of general health and prevention of cardiometabolic disease. We used 3 tools: Appraisal of Guidelines for Research and Evaluation (AGREE II), the National Academy of Medicine’s (NAM) Standards for Trustworthy Clinical Practice Guidelines, and the Frequency, Intensity, Time, and Type (FITT) score. Average AGREE II domain scores ranged from 38%−84%, and the portion of criteria fulfilled per NAM domain ranged from 7%–39%. The average FITT score for all recommendations was 2.48 out of 4. While guidelines improved according to both AGREE II and the NAM standards over time, their practicability as assessed by FITT score did not improve. Guidelines produced by governmental agencies or other nonprofit organizations, using the Grading of Recommendations Assessment, Development, and Evaluation (GRADE) approach, or fulfilling a higher number of NAM criteria tended to be of higher quality. Organizations producing physical activity guidelines can improve their quality by establishing and reporting processes for public representation, external review, and conflict of interest (COI) management. Future recommendations about physical activity should be more specific and include strategies to improve implementation. Registration no.: PROSPERO CRD42019126364. Novelty: Most physical activity recommendations are not sufficiently specific to be practically implemented. The overall quality of guidelines has improved over time, but the specificity of recommendations has not. Improved public representation, external review, and COI disclosure and management processes would improve guideline quality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".