Features, Behavioral Change Techniques, and Quality of the Most Popular Mobile Apps to Measure Physical Activity: Systematic Search in App Stores
Bibliographic record
Abstract
BACKGROUND: It is estimated that 23% of adults and 55% of older adults do not meet the recommended levels of physical activity. Thus, improving the levels of physical activity is of paramount importance, but it requires the use of low-cost resources that facilitate universal access without depleting the health system. The high number of apps available constitutes an opportunity, but it also makes it quite difficult for the layperson to select the most appropriate app. Furthermore, the information available in the app stores is often insufficient, lacks quality, and is not evidence based, and the systematic reviews fail to assess app quality using standardized and validated instruments. OBJECTIVE: The objective of this study was to systematically assess the features, content, and quality of the most popular apps that can be used to measure and, potentially, promote physical activity. METHODS: Systematic searches were conducted on Apple App Store, Google Play, and Windows Phone Store between December 2017 and January 2018. Apps were included if their primary objective was to assess the aspects of physical activity, if they had a user rating of at least 4, if their number of ratings was ≥100, and if they were free. Apps meeting these criteria were independently assessed by two reviewers regarding their general and technical information, aspects of physical activity, presence of behavioral change techniques, and quality. Data were analyzed using means and SDs or frequencies and percentages. RESULTS: Of 51 apps included, none specified the age of the target group and only one mentioned the involvement of health professionals. Most apps offered the possibility to work in background (n=50) and allowed data sharing (n=40). Regarding physical activity, most apps measured steps and distance (n=11) or steps, distance, and time (n=17). Only 18 apps, all of which measured number of steps, followed the guidelines on recommendations for physical activity. On average, 5.5 (SD 1.8) behavioral change techniques were identified per app; the most frequently used techniques were "provide feedback on performance" (n=50) and "prompt self-monitoring of behavior" (n=50). The overall quality score was 3.88 (SD 0.34). CONCLUSIONS: Although the overall quality of the apps was moderate, the quality of their content, particularly the use of international guidelines on physical activity, should be improved. Additionally, a more in-depth assessment of apps should be performed before releasing them for public use, particularly regarding their reliability and validity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".