Medical Education through Virtual Worlds: The HLTHSIM Project
Bibliographic record
Abstract
Training tools using virtual reality (VR) are becoming more popular and cost-effective to develop and are increasingly adopted; yet there is no systematic means for evaluating their usability and pedagogical effectiveness. There are a wide range of training scenarios that can be scripted, from high level simulations of emergency response systems where participants using their avatars have to make complex decisions and communicate with each other, to low-level sensor-motor skills-based trainers where surgeons can practice suturing and cutting. We propose a classification framework for simulator-based training, associating each type of simulation with a specification of the types of skills it is designed to exercise and a corresponding evaluation plan. In this framework, objective measures involving task time and error rates can be formalized at the lower levels, and related subjective and objective measures can be identified at the top. Our framework is being implemented under the auspices of a recently funded New Media project in Canada (GRAND NCE) that spans two health training and simulation facilities (CSTAR and HSERC).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".