Functionality of Top-Rated Mobile Apps for Depression: Systematic Search and Evaluation
Bibliographic record
Abstract
BACKGROUND: In the last decade, there has been a proliferation of mobile apps claiming to support the needs of people living with depression. However, it is unclear what functionality is actually provided by apps for depression, or for whom they are intended. OBJECTIVE: This paper aimed to explore the key features of top-rated apps for depression, including descriptive characteristics, functionality, and ethical concerns, to better inform the design of apps for depression. METHODS: We reviewed top-rated iPhone OS (iOS) and Android mobile apps for depression retrieved from app marketplaces in spring 2019. We applied a systematic analysis to review the selected apps, for which data were gathered from the 2 marketplaces and through direct use of the apps. We report an in-depth analysis of app functionality, namely, screening, tracking, and provision of interventions. Of the initially identified 482 apps, 29 apps met the criteria for inclusion in this review. Apps were included if they remained accessible at the moment of evaluation, were offered in mental health-relevant categories, received a review score greater than 4.0 out of 5.0 by more than 100 reviewers, and had depression as a primary target. RESULTS: The analysis revealed that a majority of apps specify the evidence base for their intervention (18/29, 62%), whereas a smaller proportion describes receiving clinical input into their design (12/29, 41%). All the selected apps are rated as suitable for children and adolescents on the marketplace, but 83% (24/29) do not provide a privacy policy consistent with their rating. The findings also show that most apps provide multiple functions. The most commonly implemented functions include provision of interventions (24/29, 83%) either as a digitalized therapeutic intervention or as support for mood expression; tracking (19/29, 66%) of moods, thoughts, or behaviors for supporting the intervention; and screening (9/29, 31%) to inform the decision to use the app and its intervention. Some apps include overtly negative content. CONCLUSIONS: Currently available top-ranked apps for depression on the major marketplaces provide diverse functionality to benefit users across a range of age groups; however, guidelines and frameworks are still needed to ensure users' privacy and safety while using them. Suggestions include clearly defining the age of the target population and explicit disclosure of the sharing of users' sensitive data with third parties. In addition, we found an opportunity for apps to better leverage digital affordances for mitigating harm, for personalizing interventions, and for tracking multimodal content. The study further demonstrated the need to consider potential risks while using depression apps, including the use of nonvalidated screening tools, tracking negative moods or thinking patterns, and exposing users to negative emotional expression content.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".