Educating residents in spine surgery: A study of Entrustable professional activities in neurosurgery and orthopedic surgery
Bibliographic record
Abstract
BACKGROUND: Surgery for spinal disorders represents some of the commonest surgical procedures performed in many countries worldwide, carried out by neurosurgeons and orthopedic surgeons. Residency training is shifting to competency-based medical education, which requires setting standards for graduating residents and their assessments. However, gaps exist in the literature regarding the parameters used for assessment and the mastery levels expected of graduating residents in the performance of common spinal procedures as defined in Entrustable Professional Activities (EPAs). The objectives of the study were to describe the assessment parameters used for residents, identify the standard of performance expected of graduating residents of EPAs of spinal procedures, and identify factors predicting the expected standard of competent performance of graduating residents. METHODS: The survey was sent to neurosurgery and orthopedic surgery Faculty requesting their recommendations on parameters of assessment and the expected standard competence performance for EPAs related to spinal procedures using our entrustment scale (A-E). RESULTS: Based on total responses, the recommended number of assessments and assessors for each EPA was 5 and 2, respectively. Regarding each specialty, there was no significant difference in the recommended number of assessments for each EPA. However, neurosurgery Faculty recommended higher number of assessors(n = 3) than orthopedic surgery Faculty(n = 2) for both posterior spinal decompression EPA(PSD) (p = 0.01) and spinal instrumentation EPA(SI) (p = 0.04). Based on total responses, 83% felt PSD was appropriate to the general practice, 86.8% considered it not too broad, and 62.3% expected entrustment level E as a graduation target. The proportions of these ratings were slightly lower for SI at 58.5%, 71.7% and 56.6%, respectively. Both specialties indicated that the EPAs were not too broad. In contrast, neurosurgery Faculty were more likely to consider these EPAs appropriate for general practice than orthopedic surgery Faculty for both PSD (94.7% vs 53.3%, p = 0.0003) and SI (68.4% vs 33.3%, p = 0.02). Moreover, neurosurgery Faculty had a higher expected standard of performance as a graduation target for both PSD (Level E 76.3% vs 26.7%, p = 0.001) and SI (Level E 65.8% vs 33.3%, p = 0.03) than orthopedic surgery Faculty. Expectations of entrustment level E for PSD was associated with the belief that the current EPA was appropriate for the general practice of their specialty with an odds ratio of 8.35 (p = 0.01, 95%CI 1.53-45.67). CONCLUSIONS: A difference exists in parameters of assessment and expected standard competence performance of spine procedures among spinal surgery specialties. In our opinion, there should be efforts to develop consensus between specialties for the sake of uniform delivery of high-quality care for patients regardless of the specialty of their surgeon. Our results will be particularly valuable to certification bodies in the assessment of spinal milestones. This study has important implications for the design of residency and fellowship education in spinal surgery internationally.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".