Effectiveness of Mobile Apps to Promote Health and Manage Disease: Systematic Review and Meta-analysis of Randomized Controlled Trials
Bibliographic record
Abstract
BACKGROUND: Interventions aimed at modifying behavior for promoting health and disease management are traditionally resource intensive and difficult to scale. Mobile health apps are being used for these purposes; however, their effects on health outcomes have been mixed. OBJECTIVE: This study aims to summarize the evidence of rigorously evaluated health-related apps on health outcomes and explore the effects of features present in studies that reported a statistically significant difference in health outcomes. METHODS: A literature search was conducted in 7 databases (MEDLINE, Scopus, PsycINFO, CINAHL, Global Index Medicus, Cochrane Central Register of Controlled Trials, and Cochrane Database of Systematic Reviews). A total of 5 reviewers independently screened and extracted the study characteristics. We used a random-effects model to calculate the pooled effect size estimates for meta-analysis. Sensitivity analysis was conducted based on follow-up time, stand-alone app interventions, level of personalization, and pilot studies. Logistic regression was used to examine the structure of app features. RESULTS: From the database searches, 8230 records were initially identified. Of these, 172 met the inclusion criteria. Studies were predominantly conducted in high-income countries (164/172, 94.3%). The majority had follow-up periods of 6 months or less (143/172, 83.1%). Over half of the interventions were delivered by a stand-alone app (106/172, 61.6%). Static/one-size-fits-all (97/172, 56.4%) was the most common level of personalization. Intervention frequency was daily or more frequent for the majority of the studies (123/172, 71.5%). A total of 156 studies involving 21,422 participants reported continuous health outcome data. The use of an app to modify behavior (either as a stand-alone or as part of a larger intervention) confers a slight/weak advantage over standard care in health interventions (standardized mean difference=0.38 [95% CI 0.31-0.45]; I2=80%), although heterogeneity was high. CONCLUSIONS: The evidence in the literature demonstrates a steady increase in the rigorous evaluation of apps aimed at modifying behavior to promote health and manage disease. Although the literature is growing, the evidence that apps can improve health outcomes is weak. This finding may reflect the need for improved methodological and evaluative approaches to the development and assessment of health care improvement apps. TRIAL REGISTRATION: PROSPERO International Prospective Register of Systematic Reviews CRD42018106868; https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=106868.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.033 | 0.086 |
| Meta-epidemiology (narrow) | 0.004 | 0.002 |
| Meta-epidemiology (broad) | 0.028 | 0.044 |
| Bibliometrics | 0.011 | 0.010 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".