Review and Evaluation of Mindfulness-Based iPhone Apps
Bibliographic record
Abstract
BACKGROUND: There is growing evidence for the positive impact of mindfulness on wellbeing. Mindfulness-based mobile apps may have potential as an alternative delivery medium for training. While there are hundreds of such apps, there is little information on their quality. OBJECTIVE: This study aimed to conduct a systematic review of mindfulness-based iPhone mobile apps and to evaluate their quality using a recently-developed expert rating scale, the Mobile Application Rating Scale (MARS). It also aimed to describe features of selected high-quality mindfulness apps. METHODS: A search for "mindfulness" was conducted in iTunes and Google Apps Marketplace. Apps that provided mindfulness training and education were included. Those containing only reminders, timers or guided meditation tracks were excluded. An expert rater reviewed and rated app quality using the MARS engagement, functionality, visual aesthetics, information quality and subjective quality subscales. A second rater provided MARS ratings on 30% of the apps for inter-rater reliability purposes. RESULTS: The "mindfulness" search identified 700 apps. However, 94 were duplicates, 6 were not accessible and 40 were not in English. Of the remaining 560, 23 apps met inclusion criteria and were reviewed. The median MARS score was 3.2 (out of 5.0), which exceeded the minimum acceptable score (3.0). The Headspace app had the highest average score (4.0), followed by Smiling Mind (3.7), iMindfulness (3.5) and Mindfulness Daily (3.5). There was a high level of inter-rater reliability between the two MARS raters. CONCLUSIONS: Though many apps claim to be mindfulness-related, most were guided meditation apps, timers, or reminders. Very few had high ratings on the MARS subscales of visual aesthetics, engagement, functionality or information quality. Little evidence is available on the efficacy of the apps in developing mindfulness.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.039 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.004 |
| Bibliometrics | 0.012 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".