Engaging Undergraduate Medical Students With Introductory Research Training via an Educational Escape Room: Mixed Methods Evaluation
Bibliographic record
Abstract
BACKGROUND: Early exposure to research methodology is essential in medical education, yet many students show limited motivation to engage with non-clinical content. Gamified strategies such as educational escape rooms (EERs) may help improve engagement, but few studies have explored their feasibility at scale or evaluated their impact beyond student satisfaction. OBJECTIVE: To assess the feasibility, engagement, and perceived educational value of a large-scale escape room specifically designed to introduce third-year medical students to the principles of diagnostic test evaluation. METHODS: We developed a low-cost immersive escape room based on a fictional diagnostic accuracy study, with six puzzles mapped to five predefined learning objectives: (1) identifying key components of a diagnostic study protocol, (2) selecting an appropriate gold-standard test, (3) defining a relevant study population, (4) building and interpreting a contingency table, and (5) critically appraising diagnostic metrics in context. The intervention was deployed to an entire class of third-year medical students across 12 sessions between March and April 2023. Each session included 60 minutes of gameplay and a 45-minute debriefing. Students completed pre-/post-intervention questionnaires assessing their knowledge of diagnostic test evaluation and perceptions of research training. Descriptive statistics and paired t-tests were used to evaluate score changes; univariate linear regressions assessed associations with demographics. Free-text comments were analyzed using Reinert's hierarchical classification. RESULTS: Among 530 participants, 490 completed the full evaluation. Many participants had limited prior exposure to escape rooms (206/490, 42% had never participated), and most reported low initial confidence with critical appraisal of scientific articles. All student teams completed the scenario, with a mean completion time of 53 (±4) minutes. Mean overall knowledge scores increased from 62/100 (±1) before to 82/100 (±2) after the activity (+32%, p<0.001). Gains were observed across all learning objectives and were not influenced by age, sex, or prior experience. Students rated the EER as highly entertaining (9.1±1.1/10) and educational (8.2±1.5/10). Following the intervention, 87% (393/452) felt more comfortable with critical appraisal of diagnostic test studies, and 79% (357/452) considered the escape room format highly appropriate for an introductory session. Thematic analysis of open-ended feedback identified six clusters, including engagement, teamwork, and perceived usefulness of the pedagogical approach. Word clouds showed a marked shift from negative to positive attitudes toward research training. CONCLUSIONS: This study demonstrates the feasibility and enthusiastic reception of a large-scale, reusable escape room aimed at teaching the fundamental principles of diagnostic test evaluation to undergraduate medical students. While not designed to cover the broader spectrum of research designs or methods, the intervention successfully addressed targeted objectives within a specific area of research appraisal. This approach may serve as a valuable entry point to engage students with evidence-based reasoning and pave the way for deeper exploration of medical research methodology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.053 | 0.040 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".