The Quality of Health Apps and Their Potential to Promote Behavior Change in Patients With a Chronic Condition or Multimorbidity: Systematic Search in App Store and Google Play
Bibliographic record
Abstract
BACKGROUND: Mobile apps offer an opportunity to improve the lifestyle of patients with chronic conditions or multimorbidity. However, for apps to be recommended in clinical practice, their quality and potential for promoting behavior change must be considered. OBJECTIVE: We aimed to investigate the quality of health apps for patients with a chronic condition or multimorbidity (defined as 2 or more chronic conditions) and their potential for promoting behavior change. METHODS: We followed the Cochrane Handbook guidelines to conduct and report this study. A systematic search of apps available in English or Danish on App Store (Apple Inc) and Google Play (Google LLC) for patients with 1 or more of the following common and disabling conditions was conducted: osteoarthritis, heart conditions (heart failure and ischemic heart disease), hypertension, type 2 diabetes mellitus, depression, and chronic obstructive pulmonary disease. For the search strategy, keywords related to these conditions were combined. One author screened the titles and content of the identified apps. Subsequently, 3 authors independently downloaded the apps onto a smartphone and assessed the quality of the apps and their potential for promoting behavior change by using the Mobile App Rating Scale (MARS; number of items: 23; score: range 0-5 [higher is better]) and the App Behavior Change Scale (ABACUS; number of items: 21; score: range 0-21 [higher is better]), respectively. We included the five highest-rated apps and the five most downloaded apps but only assessed free content for their quality and potential for promoting behavior change. RESULTS: We screened 453 apps and ultimately included 60. Of the 60 apps, 35 (58%) were available in both App Store and Google Play. The overall average quality score of the apps was 3.48 (SD 0.28) on the MARS, and their overall average score for their potential to promote behavior change was 8.07 (SD 2.30) on the ABACUS. Apps for depression and apps for patients with multimorbidity tended to have higher overall MARS and ABACUS scores, respectively. The most common app features for supporting behavior change were the self-monitoring of physiological parameters (eg, blood pressure monitoring; apps: 38/60, 63%), weight and diet (apps: 25/60, 42%), or physical activity (apps: 22/60, 37%) and stress management (apps: 22/60, 37%). Only 8 out of the 60 apps (13%) were completely free. CONCLUSIONS: Apps for patients with a chronic condition or multimorbidity appear to be of acceptable quality but have low to moderate potential for promoting behavior change. Our results provide a useful overview for patients and clinicians who would like to use apps for managing chronic conditions and indicate the need to improve health apps in terms of their quality and potential for promoting behavior change.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".