Assessing Gender Differences in Technical Skills and Confidence in Orthopaedic Surgery Residency Applicants
Bibliographic record
Abstract
INTRODUCTION: Variations in confidence for procedural skills have been demonstrated when comparing male and female medical students in surgical training. This study investigates whether differences in technical skill and self-reported confidence exist between male and female medical students applying to orthopaedic residency. METHODS: All medical students (2017 to 2020) invited to interview at a single orthopaedic residency program were prospectively evaluated on their technical skills and self-reported confidence. Objective evaluation of technical skill included scores for a suturing task as evaluated by faculty graders. Self-reported confidence in technical skills was assessed before and after completing the assigned task. Scores for male and female students were compared by age, self-identified race/ethnicity, number of publications at the time of application, athletic background, and US Medical Licensing Examination Step 1 score. RESULTS: Two hundred sixteen medical students were interviewed, of which 73% were male (n = 158). No gender differences were observed in suture task technical skill scores or mean difference in simultaneous visual task scores. The mean change from pre-task and post-task self-reported confidence scores was similar between sexes. Although female students trended toward lower post-task self-reported confidence scores compared with male students, this did not achieve statistical significance. Lower self-reported confidence was associated with a higher US Medical Licensing Examination score and with attending a private medical school. DISCUSSION: No difference in technical skill or confidence was found between male and female applicants to a single orthopaedic surgery residency program. Female applicants trended toward self-reporting lower confidence than male applicants in post-task evaluations. Differences in confidence have been shown previously in surgical trainees, which may suggest that differences in skill and confidence may develop during residency training.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".