Review of Cognitive Biases in ACGME Milestones Training Assessments in Post-graduate Medical Education Programs
Bibliographic record
Abstract
The course of study for young physicians for post-graduate training is an exciting and life-changing opportunity, one that is filled with the relentless optimism of intellectual discovery and personal growth and development. The American Council for Graduate Medical Education (ACGME) is a non-profit private council that evaluates and accredits internship, residency, and fellowship programs. The role of the ACGME is to oversee curriculums, training environments, and specialty evaluation standards to ensure satisfactory competency leading to board eligibility and certification in the respected field of study. The ACGME has the monumental task of guiding educational standards that are designed to both protect the public welfare and further educational programs. Many educational standards are objective, such as quantitative performance on examinations, involvement in research, and involvement in systems development and quality improvement. However, key clinical performance measures are based on prior training and experience. Over the last several years, studies examining rates of abuse and discrimination during post-graduate medical training in both the United States and Canadian studies, which have reported alarmingly high rates of 50%. With the increasing utility and availability of social media, such issues have become more transparent to the public. A plethora of studies has been conducted, examining physician biases towards patients, practice changes, insurance company regulations, and evolving healthcare systems. However, a significant amount of evaluation is merited when examining individual institutional cultures and the educational environments that harbor them. We wish to examine the role of ever-evolving specialty-specific ACGME-instituted educational milestones in Internal Medicine and Opthalmology in the context of potential cognitive biases and their implementation within post-graduate training programs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.021 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".