Developing and Evaluating Learning Theory-Informed Extended Reality Simulations in Medical Education: Lessons from a Pilot Randomised Controlled Trial
Bibliographic record
Abstract
Introduction: Despite hopes that the adoption of extended reality (XR) technology in medical simulation could offer a scalable and accessible teaching option in medical education, the development of novel XR-enhanced simulation training interventions is rarely underpinned by a learning theoretical framework. Additionally, there is a notable lack of robust experimental studies evaluating the effectiveness of XR-enhanced simulation in comparison to more established simulation modalities, with existing and validated assessment instruments seldom used to assess outcomes. Addressing these gaps, I worked with an interdisciplinary team (University of Cambridge, Cambridge University Hospitals, and GigXR) to develop and evaluate ‘Holoscenarios,’ a mixed reality (MR) enhanced simulation training intervention grounded in constructivism, designed for medical students to practice the assessment and management of acute medical scenarios. I then designed and implemented a pilot randomised controlled trial (RCT) to evaluate the effectiveness of Holoscenarios versus manikin-based simulation (MBS). The primary objective of this study was to assess the processes, resources, and management strategies required for running an experimental study comparing the effectiveness of an XR-enhanced simulation intervention with an established simulation training modality. Additionally, I aimed to gather preliminary data on whether Holoscenarios was non-inferior to MBS in providing medical students with the technical (TS) and non-technical skills (NTS) required to manage an acutely deteriorating patient. Methodology: The completed module of Holoscenarios comprised three 20-minute interactive acute medical scenarios surrounding the assessment and management of an acutely deteriorating patient. Scenarios were depicted using holographic overlays viewed through the HoloLens 2, and aligned with learning outcomes based on the General Medical Council’s (GMC) Outcomes for Graduates. A pilot RCT was integrated into final-year medical students' simulation training days, utilising a pre-post-test control group design. Over seven training days, with six students attending each day, participants were randomly allocated to an MBS or MR simulation training course. Participants in each course completed three facilitated simulated acute medical scenarios matched in content. Participants' technical skills (TS) and non-technical skills (NTS) surrounding the assessment and management of an acutely deteriorating patient were evaluated at baseline and immediately after training courses via a 10-minute pre-test and post-test, each comprising an observed simulated acute medical scenario. All pre- and post-tests were assessed using pre-existing standardised assessment instruments: the Queens’ Simulation Assessment Tool (QSAT), a modifiable anchored rating scale designed for the competency-based assessment of simulated resuscitation scenarios, and the Ottawa Global Rating Scale (GRS), a behavioural rating scale designed to assess six categories of NTS in simulated emergency scenarios, both of which aligned with the learning outcome of Holoscenarios. To measure students’ engagement and experience following training, the Satisfaction with Simulation Experience Scale (SSE) was employed. Additionally, self-reported confidence was evaluated at baseline and post-training utilising a newly developed Likert scale that rated participants' self-reported confidence in performing skills related to assessing and managing an acutely deteriorating patient. The scores from these assessment rubrics were then compared from pre-test to post-test within groups utilising a Wilcoxon Signed-Rank Test, with changes in scores from pre-training to post-training compared between the MR and MBS groups using a Mann-Whitney U test. Results: Data collection was completed in May 2024, with 28 participants completing the pilot RCT. Preliminary findings showed that both MR and MBS groups showed a significant improvement in TS and NTS across all assessment domains from pre-training to post-training. Notably, there were no significant differences in the improvements between the two theoretically based interventions, suggesting that MR is comparable to traditional MBS training methods in improving the TS and NTS required to assess and manage an acutely deteriorating patient. Moreover, students in both groups reported increased self-reported confidence in applying the skills required to assess and manage acutely deteriorating patients. While the educational outcomes were promising, I encountered significant operational challenges whilst implementing this RCT. This study necessitated integration into final-year medical students’ existing training days to mitigate simulation centre costs. Data collection periods were therefore limited to students’ existing rotas, restricting the number of participants available to enrol in the study. Both simulation training courses required skilled staff for setup and operation, with MR courses requiring additional technical support. Maintaining high standards for both MBS and MR courses necessitated three simulation technicians, who underwent 2 days of training before the simulation courses. Three clinical facilitators per simulation day were required, relying on clinical staff to take time away from clinical commitments voluntarily. Similarly, maintaining standardisation of the intervention and assessment between groups was compromised due to differences in facilitation, group composition and variability between assessors. Discussion: This project sought to address the existing gap in theoretical frameworks underpinning the development of innovative XR-enhanced simulation training interventions. Grounded in constructivist learning theory, Holoscenarios emphasises a strategy for leveraging MR technology to create immersive and realistic scenarios that promote problem-solving and reflection on decision-making. Furthermore, this study provides preliminary evidence suggesting that MR-enhanced simulation is at least comparable to MBS in improving the TS, NTS, and self-reported confidence related to the assessment and management of acutely deteriorating patients among undergraduate medical students. These improvements align with the constructivist principles underpinning the design of Holoscenarios, which suggests that an interactive and immersive MR experience enables students to engage meaningfully with clinical problems to build upon existing knowledge and develop new skills. Furthermore, this research addresses the lack of robust experimental research supporting the use of XR-enhanced simulation in medical education. Initial findings contribute to the existing literature by highlighting the unique operational and logistical challenges of running experimental research in this domain, which are commonly encountered in medical education and simulation research. These challenges were analysed to formulate future recommendations for mitigating these challenges whilst identifying the resources required to do so, and acknowledging issues that may be unsurmountable in this context. In doing so, this study lays a foundational framework for executing future large-scale experimental research evaluating an XR-enhanced simulation intervention. As the development of XR-enhanced teaching tools in medical education evolves, future research designs must account for the challenges and mitigation strategies outlined in this work.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".