Analysis of Orthopaedic In-Training Examination Trauma Questions: 2017 to 2021
Bibliographic record
Abstract
INTRODUCTION: The Orthopaedic In-Training Examination (OITE) is a multiple-choice examination developed by the American Academy of Orthopaedic Surgeons annually since 1963 to assess orthopaedic residents' knowledge. This study's purpose is to analyze the 2017 to 2021 OITE trauma questions to aid orthopaedic residents preparing for the examination. METHODS: The 2017 to 2021 OITEs on American Academy of Orthopaedic Surgeons' ResStudy were retrospectively reviewed to identify trauma questions. Question topic, references, and images were analyzed. Two independent reviewers classified each question by taxonomy. RESULTS: Trauma represented 16.6% (204/1,229) of OITE questions. Forty-nine percent of trauma questions included images (100/204), 87.0% (87/100) of which contained radiographs. Each question averaged 2.4 references, of which 94.9% were peer-reviewed articles and 46.8% were published within 5 years of the respective OITE. The most common taxonomic classification was T1 (46.1%), followed by T3 (37.7%) and T2 (16.2%). DISCUSSION: Trauma represents a notable portion of the OITE. Prior OITE trauma analyses were published greater than 10 years ago. Since then, there has been an increase in questions with images and requiring higher cognitive processing. The Journal of Orthopaedic Trauma (24.7%), Journal of the American Academy of Orthopaedic Surgeons (10.1%), and Journal of Bone and Joint Surgery, American Volume (9.3%) remain the most cited sources.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.007 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".