Digital Mental Health Coaching in Clinically Diverse Populations: Controlled Engagement and Outcomes Study
Bibliographic record
Abstract
Background: Digital coaching programs, offering virtually delivered mental health care by coaches and companion apps, are an increasingly popular care model designed to increase accessibility and reduce strain on traditional mental health care systems. Initial studies suggest these programs can produce a range of positive mental health outcomes; however, methodological limitations and a focus on homogeneous, subclinical populations have constrained conclusions about their effectiveness, especially in diverse and clinically severe samples. Objective: This study aimed to evaluate the impact of an evidence-based digital mental health coaching program in a clinically and demographically diverse sample. The study compared engagement with app-based content and changes in depression, anxiety, and stress symptoms, over the course of 1 month, among users who received coaching versus those who used the app alone (controls). Methods: Program users (N=64) were categorized as coaching users (attending at least 1 session) or controls (app-only users). Depression, anxiety, and stress symptoms were assessed using the Depression Anxiety Stress Scale-21 at baseline and after 30 days. Engagement with app content was also measured. Between-group differences were analyzed using t tests and mixed multivariate analysis of covariance models, with follow-up sensitivity analysis of covariance analyses (controlling for age). Results: Participants were diverse in terms of demographics and clinical severity, with half reporting severe to extremely severe depression and nearly half reporting severe to extremely severe anxiety or stress at baseline. A repeated-measures multivariate analysis of covariance revealed a significant group-by-time interaction (P=.02), indicating greater symptom reduction among coaching users, primarily driven by changes in anxiety and stress. Follow-up analyses of covariance exploring symptom-specific patterns, excluding participants with subclinical baseline symptoms, yielded significant group-by-time interactions across depression (P=.04), anxiety (P=.003), and stress (P=.03). Engagement with app-based content did not significantly differ between the groups (P=.20), suggesting coaching's effectiveness was not contingent on differential app usage. Conclusions: This study demonstrates that digital mental health coaching can significantly improve clinical outcomes, even in diverse and clinically severe populations. These findings challenge the notion that coaching is only effective for subclinical or high-functioning individuals and highlight its potential to extend the reach of mental health care to underserved communities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".