Current Landscape and Future Directions for Mental Health Conversational Agents for Youth: Scoping Review
Bibliographic record
Abstract
BACKGROUND: Conversational agents (CAs; chatbots) are systems with the ability to interact with users using natural human dialogue. They are increasingly used to support interactive knowledge discovery of sensitive topics such as mental health topics. While much of the research on CAs for mental health has focused on adult populations, the insights from such research may not apply to CAs for youth. OBJECTIVE: This study aimed to comprehensively evaluate the state-of-the-art research on mental health CAs for youth. METHODS: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, we identified 39 peer-reviewed studies specific to mental health CAs designed for youth across 4 databases, including ProQuest, Scopus, Web of Science, and PubMed. We conducted a scoping review of the literature to evaluate the characteristics of research on mental health CAs designed for youth, the design and computational considerations of mental health CAs for youth, and the evaluation outcomes reported in the research on mental health CAs for youth. RESULTS: We found that many mental health CAs (11/39, 28%) were designed as older peers to provide therapeutic or educational content to promote youth mental well-being. All CAs were designed based on expert knowledge, with a few that incorporated inputs from youth. The technical maturity of CAs was in its infancy, focusing on building prototypes with rule-based models to deliver prewritten content, with limited safety features to respond to imminent risk. Research findings suggest that while youth appreciate the 24/7 availability of friendly or empathetic conversation on sensitive topics with CAs, they found the content provided by CAs to be limited. Finally, we found that most (35/39, 90%) of the reviewed studies did not address the ethical aspects of mental health CAs, while youth were concerned about the privacy and confidentiality of their sensitive conversation data. CONCLUSIONS: Our study highlights the need for researchers to continue to work together to align evidence-based research on mental health CAs for youth with lessons learned on how to best deliver these technologies to youth. Our review brings to light mental health CAs needing further development and evaluation. The new trend of large language model-based CAs can make such technologies more feasible. However, the privacy and safety of the systems should be prioritized. Although preliminary evidence shows positive trends in mental health CAs, long-term evaluative research with larger sample sizes and robust research designs is needed to validate their efficacy. More importantly, collaboration between youth and clinical experts is essential from the early design stages through to the final evaluation to develop safe, effective, and youth-centered mental health chatbots. Finally, best practices for risk mitigation and ethical development of CAs with and for youth are needed to promote their mental well-being.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".