GEMINI MedED: Leveraging ‘Big Data’ to Understand Clinical Practice Variation in Resident Physicians
Bibliographic record
Abstract
Introduction: Residency training is an important source of clinical practice variation, given the known association between the training environment and future patient outcomes. However, current assessment approaches in postgraduate medical education do not routinely measure practice variation during training, leaving a gap in understanding how residency education could be optimized to improve patient care. Objectives: This retrospective cohort study aimed to measure and interpret clinical practice variation between senior internal medicine (IM) residents. First, quality improvement and clinical practice guidelines were used to select candidate resident-sensitive quality measures (RSQMs), defined as measures which were meaningful to patient care and mostly attributable to residents. Second, RSQMs were used to measure and interpret potential causes of practice variation. Methods: This study used electronic health record (EHR) data from the General Medicine Inpatient Initiative Medical Education Database (GEMINI MedED), which links 793 senior IM residents at the University of Toronto to clinical data from 132,291 patients between 2010-2019. Ten RSQMs were selected related to pneumonia-specific and general IM care, then categorized based on the expected patterns of variation. Descriptive statistics were used to characterize variation in each RSQM. Results: Nine of ten candidate RSQMs were ultimately included. First line antibiotic ordering in pneumonia was performed for only 52.6% (7,085 / 13,470) of patients with wide variation. Wide variation in performance was observed for metrics categorized as discretionary (context-dependent). Potentially inappropriate transfusions were performed for 0.02% (26 / 132,291) of patients with low variation. Discussion: This study of 793 senior IM residents who cared for 132,291 newly admitted patients over 10 academic years demonstrated wide variation in clinical practice for pneumonia and general IM care. By increasing our understanding of how to characterize resident clinical practice variation, these findings offer several conceptual insights: (1) Considering the educational potential of identified variation, (2) A broader conceptualization of attribution as intentionally titratable and relevant to both diagnoses and patients, and (3) The need for a concept-driven approach to guide the use of EHR data. Overall, this study represents an incremental step towards better alignment of feedback, learning, and assessment in residency education with improved patient outcomes. ---------- ERRATUM 2026-05-13 This manuscript characterized RSQM data using the mean, standard deviation, and interquartile range (IQR) as key descriptive statistics (Table 7). However, because most variables were skewed, the subsequently peer-reviewed published manuscript reported medians and IQRs instead. Additionally, after completion of this thesis manuscript, an error was identified in the initial query used to capture second-line antibiotic orders. Correcting this error resulted in the identification of additional second-line antibiotic orders. Although the conceptual findings and overall conclusions of this manuscript remain unchanged, readers are encouraged to consult the final peer-reviewed publication for the definitive quantitative results of this study: https://jamanetwork.com/journals/jamanetworkopen/fullarticle/2848775
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.070 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".