Extended Reality–Enhanced Mental Health Consultation Training: Quantitative Evaluation Study
Bibliographic record
Abstract
BACKGROUND: The use of extended reality (XR) technologies in health care can potentially address some of the significant resource and time constraints related to delivering training for health care professionals. While substantial progress in realizing this potential has been made across several domains, including surgery, anatomy, and rehabilitation, the implementation of XR in mental health training, where nuanced humanistic interactions are central, has lagged. OBJECTIVE: Given the growing societal and health care service need for trained mental health and care workers, coupled with the heterogeneity of exposure during training and the shortage of placement opportunities, we explored the feasibility and utility of a novel XR tool for mental health consultation training. Specifically, we set out to evaluate a training simulation created through collaboration among software developers, clinicians, and learning technologists, in which users interact with a virtual patient, "Stacey," through a virtual reality or augmented reality head-mounted display. The tool was designed to provide trainee health care professionals with an immersive experience of a consultation with a patient presenting with perinatal mental health symptoms. Users verbally interacted with the patient, and a human instructor selected responses from a repository of prerecorded voice-acted clips. METHODS: In a pilot experiment, we confirmed the face validity and usability of this platform for perinatal and primary care training with subject-matter experts. In our follow-up experiment, we delivered personalized 1-hour training sessions to 123 participants, comprising mental health nursing trainees, general practitioner doctors in training, and students in psychology and medicine. This phase involved a comprehensive evaluation focusing on usability, validity, and both cognitive and affective learning outcomes. RESULTS: We found significant enhancements in learning metrics across all participant groups. Notably, there was a marked increase in understanding (P<.001) and motivation (P<.001), coupled with decreased anxiety related to mental health consultations (P<.001). There were also significant improvements to considerations toward careers in perinatal mental health (P<.001). CONCLUSIONS: Our findings show, for the first time, that XR can be used to provide an effective, standardized, and reproducible tool for trainees to develop their mental health consultation skills. We suggest that XR could provide a solution to overcoming the current resource challenges associated with equipping current and future health care professionals, which are likely to be exacerbated by workforce expansion plans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".