Utilizing Grasp Monitoring to Predict Microsurgical Expertise
Bibliographic record
Abstract
INTRODUCTION: Most microsurgical procedures require the surgeon to use tools to grasp and hold fragile objects in the surgical site. Prior research on grasping in surgery has mostly either been in other surgical techniques or used grasping as an auxiliary metric. We focus on microsurgery and investigate what grasping can tell about microsurgical skill and suturing performance. This study lays groundwork for using automatic detection of grasps to evaluate surgical skill. METHODS: Five expert surgeons and six novices completed sutures on a microsurgical training board. Video recordings of the performance were annotated for the number of grasps, while an eye tracker recorded the participants' pupil dilations for cognitive workload assessment. Performance was measured with suturing duration and the University of Western Ontario Microsurgical Skills Assessment instrument (UWOMSA). Differences in skill, suturing performance and cognitive workload were compared with grasping behavior. RESULTS: Novices needed significantly more grasps to complete sutures and failed to grasp more often than the experts. The number of grasps affected the suturing duration more in novices. Decreasing suturing efficiency as measured by UWOMSA instrument was associated with increase in grasps, even when we controlled for overall skill differences. Novices displayed larger pupil dilations when averaged over a sufficiently large sample, and the difference increased after the grasp. CONCLUSIONS: Grasping action during microsurgical procedures can be used as a conceptually simple yet objective proxy in microsurgical performance assessment. If the grasps could be detected automatically, they could be used to aid in computational evaluation of surgical trainees' performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".