How effective are physical activity interventions when they are scaled-up: a systematic review
Bibliographic record
Abstract
BACKGROUND: The 'scale-up' of effective physical activity interventions is required if they are to yield improvements in population health. The purpose of this study was to systematically review the effectiveness of community-based physical activity interventions that have been scaled-up. We also sought to explore differences in the effect size of these interventions compared with prior evaluations of their efficacy in more controlled contexts, and describe adaptations that were made to interventions as part of the scale-up process. METHODS: We performed a search of empirical research using six electronic databases, hand searched reference lists and contacted field experts. An intervention was considered 'scaled-up' if it had been intentionally delivered on a larger scale (to a greater number of participants, new populations, and/or by means of different delivery systems) than a preceding randomised control trial ('pre-scale') in which a significant intervention effect (p < 0.05) was reported on any measure of physical activity. Effect size differences between pre-scale and scaled up interventions were quantified ([the effect size reported in the scaled-up study / the effect size reported in the pre-scale-up efficacy trial] × 100) to explore any scale-up 'penalties' in intervention effects. RESULTS: We identified 10 eligible studies. Six scaled-up interventions appeared to achieve significant improvement on at least one measure of physical activity. Six studies included measures of physical activity that were common between pre-scale and scaled-up trials enabling the calculation of an effect size difference (and potential scale-up penalty). Differences in effect size ranged from 132 to 25% (median = 58.8%), suggesting that most scaled-up interventions typically achieve less than 60% of their pre-scale effect size. A variety of adaptations were made for scale-up - the most common being mode of delivery. CONCLUSION: The majority of interventions remained effective when delivered at-scale however their effects were markedly lower than reported in pre-scale trials. Adaptations of interventions were common and may have impacted on the effectiveness of interventions delivered at scale. These outcomes provide valuable insight for researchers and public health practitioners interested in the design and scale-up of physical activity interventions, and contribute to the growing evidence base for delivering health promotion interventions at-scale. TRIAL REGISTRATION: PROSPERO CRD42020144842 .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.006 | 0.004 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".