Quality of Physical Activity Apps: Systematic Search in App Stores and Content Analysis
Bibliographic record
Abstract
BACKGROUND: Physical inactivity is a major contributor to the development and persistence of chronic diseases. Mobile health apps that foster physical activity have the potential to assist in behavior change. However, the quality of the mobile health apps available in app stores is hard to assess for making informed decisions by end users and health care providers. OBJECTIVE: This study aimed at systematically reviewing and analyzing the content and quality of physical activity apps available in the 2 major app stores (Google Play and App Store) by using the German version of the Mobile App Rating Scale (MARS-G). Moreover, the privacy and security measures were assessed. METHODS: A web crawler was used to systematically search for apps promoting physical activity in the Google Play store and App Store. Two independent raters used the MARS-G to assess app quality. Further, app characteristics, content and functions, and privacy and security measures were assessed. The correlation between user star ratings and MARS was calculated. Exploratory regression analysis was conducted to determine relevant predictors for the overall quality of physical activity apps. RESULTS: Of the 2231 identified apps, 312 met the inclusion criteria. The results indicated that the overall quality was moderate (mean 3.60 [SD 0.59], range 1-4.75). The scores of the subscales, that is, information (mean 3.24 [SD 0.56], range 1.17-4.4), engagement (mean 3.19 [SD 0.82], range 1.2-5), aesthetics (mean 3.65 [SD 0.79], range 1-5), and functionality (mean 4.35 [SD 0.58], range 1.88-5) were obtained. An efficacy study could not be identified for any of the included apps. The features of data security and privacy were mainly not applied. Average user ratings showed significant small correlations with the MARS ratings (r=0.22, 95% CI 0.08-0.35; P<.001). The amount of content and number of functions were predictive of the overall quality of these physical activity apps, whereas app store and price were not. CONCLUSIONS: Apps for physical activity showed a broad range of quality ratings, with moderate overall quality ratings. Given the present privacy, security, and evidence concerns inherent to most rated apps, their medical use is questionable. There is a need for open-source databases of expert quality ratings to foster informed health care decisions by users and health care providers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.086 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.045 | 0.023 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".