What is the Current State of Extended Reality Use in Otolaryngology Training? A Scoping Review
Bibliographic record
Abstract
OBJECTIVE: To map current literature on the educational use of extended reality (XR) in Otolaryngology-Head and Neck Surgery (OHNS) to inform teaching and research. STUDY DESIGN: Scoping Review. METHODS: A scoping review was conducted, identifying literature through MEDLINE, Ovid Embase, and Web of Science databases. Findings were reported according to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for scoping review checklist. Studies were included if they involved OHNS trainees or medical students who used XR for an educational purpose in OHNS. XR was defined as: fully-immersive virtual reality (VR) using head-mounted displays (HMDs), non-immersive and semi-immersive VR, augmented reality (AR), or mixed reality (MR). Data on device use were extracted, and educational outcomes were analyzed according to Kirkpatrick's evaluation framework. RESULTS: Of the 1,434 unique abstracts identified, 40 articles were included. All articles reported on VR; none discussed AR or MR. Twenty-nine articles were categorized as semi-immersive, none used occlusive HMDs therefore, none met modern definitions of immersive VR. Most studies (29 of 40) targeted temporal bone surgery. Using the Kirkpatrick four-level evaluation model, all studies were limited to level-1 (learner reaction) or level-2 (knowledge or skill performance). CONCLUSIONS: Current educational applications of XR in OHNS are limited to VR, do not fully immerse participants and do not assess higher-level learning outcomes. The educational OHNS community would benefit from a shared definition for VR technology, assessment of skills transfer (level-3 and higher), and deliberate testing of AR, MR, and procedures beyond temporal bone surgery. Laryngoscope, 133:227-234, 2023.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".