Apps With Maps—Anxiety and Depression Mobile Apps With Evidence-Based Frameworks: Systematic Search of Major App Stores
Bibliographic record
Abstract
BACKGROUND: Mobile mental health apps have become ubiquitous tools to assist people in managing symptoms of anxiety and depression. However, due to the lack of research and expert input that has accompanied the development of most apps, concerns have been raised by clinicians, researchers, and government authorities about their efficacy. OBJECTIVE: This review aimed to estimate the proportion of mental health apps offering comprehensive therapeutic treatments for anxiety and/or depression available in the app stores that have been developed using evidence-based frameworks. It also aimed to estimate the proportions of specific frameworks being used in an effort to understand which frameworks are having the most influence on app developers in this area. METHODS: A systematic review of the Apple App Store and Google Play store was performed to identify apps offering comprehensive therapeutic interventions that targeted anxiety and/or depression. The PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) checklist was adapted to guide this approach. RESULTS: Of the 293 apps shortlisted as offering a therapeutic treatment for anxiety and/or depression, 162 (55.3%) mentioned an evidence-based framework in their app store descriptions. Of the 293 apps, 88 (30.0%) claimed to use cognitive behavioral therapy techniques, 46 (15.7%) claimed to use mindfulness, 27 (9.2%) claimed to use positive psychology, 10 (3.4%) claimed to use dialectical behavior therapy, 5 (1.7%) claimed to use acceptance and commitment therapy, and 20 (6.8%) claimed to use other techniques. Of the 162 apps that claimed to use a theoretical framework, only 10 (6.2%) had published evidence for their efficacy. CONCLUSIONS: The current proportion of apps developed using evidence-based frameworks is unacceptably low, and those without tested frameworks may be ineffective, or worse, pose a risk of harm to users. Future research should establish what other factors work in conjunction with evidence-based frameworks to produce efficacious mental health apps.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".