Hybrid Augmented Reality Simulator: Preliminary Construct Validation of Laparoscopic Smoothness in a Urology Residency Program
Bibliographic record
Abstract
PURPOSE: We examined the usefulness, reliability and applicability of the smoothness metric of the ProMIS hybrid simulator (Haptica, Dublin, Ireland) for a urology residency program. MATERIALS AND METHODS: A total of 15 urology residents divided into junior and senior cohorts were followed prospectively for 6 training sessions. Validated McGill Inanimate System for Training and Evaluation of Laparoscopic Skills (MISTELS) laparoscopic tasks were used. The ProMIS hybrid simulator smoothness parameter, a unit-free metric of movement efficiency, was recorded using 3-dimensional visual tracking technology. Results were compared between cohorts at the midpoint and end of the defined training sessions. End of study junior means were also retrospectively compared to senior mid training means. Statistical significance was determined using the Mann-Whitney U test (alpha = 0.05). RESULTS: Statistically significant differences between 8 junior and 7 senior cohorts were measured in all MISTELS tasks. A statistically significant performance variation was also detected at the mid and end testing times. When juniors and seniors were compared between sessions 1 and 3, and 4 and 6, statistically significant performance improvements were noted. Lastly, statistical differences were also maintained when mid session senior means were compared to end of session junior means. A 38% improvement in task completion in the senior cohort as well as a 10-fold decrease in variance was observed compared to a 12% improvement in juniors, indicating greater efficiency of movement in seniors. CONCLUSIONS: The laparoscopic smoothness metric in the hybrid simulator demonstrated construct validity by effectively differentiating between experienced and novice urology residents using validated MISTELS tasks. The outcome suggests that the hybrid simulator smoothness metric is a valuable asset in residency programs for preparatory training for live operative experience, allowing improved trainee assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".