Stages of Milestones Implementation: A Template Analysis of 16 Programs Across 4 Specialties
Bibliographic record
Abstract
BACKGROUND: Since 2013, US residency programs have used the competency-based framework of the Milestones to report resident progress and to provide feedback to residents. The implementation of Milestones-based assessments, clinical competency committee (CCC) meetings, and processes for providing feedback varies among programs and warrants systematic examination across specialties. OBJECTIVE: We sought to determine how varying assessment, CCC, and feedback implementation strategies result in different outcomes in resource expenditure and stakeholder engagement, and to explore the contextual forces that moderate these outcomes. METHODS: From 2017 to 2018, interviews were conducted of program directors, CCC chairs, and residents in emergency medicine (EM), internal medicine (IM), pediatrics, and family medicine (FM), querying their experiences with Milestone processes in their respective programs. Interview transcripts were coded using template analysis, with the initial template derived from previous research. The research team conducted iterative consensus meetings to ensure that the evolving template accurately represented phenomena described by interviewees. RESULTS: Forty-four individuals were interviewed across 16 programs (5 EM, 4 IM, 5 pediatrics, 3 FM). We identified 3 stages of Milestone-process implementation, including a resource-intensive early stage, an increasingly efficient transition stage, and a final stage for fine-tuning. CONCLUSIONS: Residency program leaders can use these findings to place their programs along an implementation continuum and gain an understanding of the strategies that have enabled their peers to progress to improved efficiency and increased resident and faculty engagement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".