Selection for family medicine residency training in Canada: How consistently are the same students ranked by different programs?
Bibliographic record
Abstract
OBJECTIVE: To examine the consistency of the ranking of Canadian and US medical graduates who applied to Canadian family medicine (FM) residency programs between 2007 and 2013. DESIGN: Descriptive cross-sectional study. SETTING: Family medicine residency programs in Canada. PARTICIPANTS: All 17 Canadian medical schools allowed access to their anonymized program rank-order lists of students applying to FM residency programs submitted to the first iteration of the Canadian Resident Matching Service match from 2007 to 2013. MAIN OUTCOME MEASURES: The rank position of medical students who applied to more than 1 FM residency program on the rank-order lists submitted by the programs. Anonymized ranking data submitted to the Canadian Resident Matching Service from 2007 to 2013 by all 17 FM residency programs were used. Ranking data of eligible Canadian and US medical graduates were analyzed to assess the within-student and between-student variability in rank score. These covariance parameters were then used to calculate the intraclass correlation coefficient (ICC) for all programs. Program descriptions and selection criteria were also reviewed to identify sites with similar profiles for subset ICC analysis. RESULTS: Between 2007 and 2013, the consistency of ranking by all programs was fair at best (ICC = 0.34 to 0.39). The consistency of ranking by larger urban-based sites was weak to fair (ICC = 0.23 to 0.36), and the consistency of ranking by sites focusing on training for rural practice was weak to moderate (ICC = 0.16 to 0.55). CONCLUSION: In most cases, there is a low level of consistency of ranking of students applying for FM training in Canada. This raises concerns regarding fairness, particularly in relation to expectations around equity and distributive justice in selection processes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".