Development and validation of a virtual reality transrectal ultrasound guided prostatic biopsy simulator
Bibliographic record
Abstract
Surgical training has undergone some remarkable changes over the past 2 decades. No longer is the operating room the sole training ground for medical students and residents learning surgery. Basic skills, such as knot tying and laparoscopic suturing, can be practiced in the confines of a surgical skills laboratory.1 Studies have shown that ex-vivo surgical skills training can lead to significant improvement in actual intra-operative performance.2 Models used to teach learners have varied in fidelity, in terms of how realistic the model looks, from simple bench models to complex virtual reality (VR) model. However, a high fidelity or virtual reality model is not synonymous with an excellent teaching tool.3 The authors of this study have developed a VR model that captures the critical constructs of performing a transrectal ultrasound and biopsy (TRUS-BX).4 Targeting is the foundation of TRUS-BX and teaching a learner how to target is the focus of this simulator. The VR simulator developed by the University of Western Ontario group incorporates real patient 3D TRUS data to train and test a learner’s ability to target 12 virtual targets. The simulator calculates the accuracy of each of the biopsy taken and also records time. Face and content validity, through self-made questionnaire, showed this model to have a very realistic feel and simulation of a TRUS-BX. The VR TRUS-BX simulator was also able to discriminate performance between experts and novices. Furthermore, improvements were demonstrated while using this model. As the authors surmized, this VR simulator shows promise as a tool that can be used to train residents and serve a role in continuing professional development. We have seen great progress being made in surgical education research and advance in the science of technical skills assessment. However, there is also much more to be desired. As technology advances and increasingly more high fidelity VR simulators become available, there is a need for standardized methods in evaluating these new simulators. Measuring a simulators face and content validity needs to be more than experts’ opinion that “yes-this looks and feels like the real thing.” A stringent, reliable and accurate tool to measure face and content validity would allow meaningful comparison between simulator offerings from different companies. Measuring construct validity using subjects of varying experience has become a standard and acceptable method of simulator validation in the surgical education field. What would be highly desirable is a simulator with high predictive validity. Can the performance in the simulator predict the performance in the real-world? This would have significant implications on residency training and would introduce a high-stakes technical skills examination using simulator to attest to one’s competence. To achieve this goal, one needs to be able to measure technical performance in the operating room, which still remains the “holy grail” of surgical education. Ethical issues of live patients and the lack of unobtrusive and practical methods of intra-operative assessment remain impeding factors. Surgical education research remains a field that is evolving and it is encouraging to see well-thought simulators being designed and evaluated in a thorough manner.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".