Evaluating surgeons’ informed decision making skills: pilot test using a videoconferenced standardised patient
Bibliographic record
Abstract
BACKGROUND: Standardised patients (SPs) are effective in evaluating communication skills, but not every training site may have the resources to develop and maintain SP programmes. OBJECTIVES: To test whether videoconferencing technology (VT) could enable an interaction between an SP and an orthopaedic surgeon that would allow the SP to accurately evaluate the surgeon's informed decision making (IDM) skills. We also assessed whether this sort of interaction was acceptable to orthopaedic surgeons as a means of learning IDM skills. METHODS: We trained an SP to represent a 75-year-old woman considering hip replacement surgery. Orthopaedic surgeons in Chicago individually consulted with the SP in Philadelphia; each participant could see and hear the other on large television screens. The SP evaluated the surgeons' advice using a 23-item checklist of IDM elements, and gave each surgeon verbal and written feedback on his IDM skills. The surgeons then gave their evaluations of the exercise. RESULTS: Twenty-two surgeons completed the project. The SP was > or = 80% accurate in classifying 20 of the 23 IDM skills when compared to a clinician rater. Although 12 (55%) of the orthopaedic surgeons felt that some aspects of the technology were distracting, most were pleased with it, and 19 of 22 (86%) would recommend the videoconferenced SP interaction to their colleagues as a means of learning IDM skills. CONCLUSIONS: These results suggest that VT allows accurate evaluation of IDM skills in a format that is acceptable to orthopaedic surgeons. Videoconferencing technology may be useful in long-distance SP communication assessment for a variety of learners.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.109 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".