The Relationship Between Accreditation Cycle and Licensing Examination Scores: A National Look
Bibliographic record
Abstract
PURPOSE: Accreditation aims to ensure all training programs meet agreed-upon standards of quality. The process is complex, resource intensive, and costly. Its benefits are difficult to assess because contextual confounds obscure comparisons between systems that do and do not include accreditation. This study explores accreditation's influence "within system" by investigating the relationship between accreditation cycle and performance on a national licensing examination. METHOD: Scores on the computer-based portion of the Medical Council of Canada Qualifying Examination Part I, from 1993 to 2017, were examined for all 17 Canadian medical schools. Typically completed upon graduation from medical school, results within each year were transformed for comparability across administrations and linked to timing within each school's accreditation cycle. ANOVAs were used to assess the relationship between accreditation timing and examination scores. Secondary analyses isolated 4-year from 3-year training programs and separated data generated before versus after implementation of a national midcycle informal review program. RESULTS: Performance on the licensing exam was highest during and shortly after an accreditation site visit, falling significantly until the midpoint in the accreditation cycle (d = 0.47) before rising again. This pattern disappeared after introduction of informal interim review, but too little data have accumulated post implementation to determine if interim review is sufficient to break the influence of accreditation cycle. CONCLUSIONS: Formal, externally driven, accreditation cycles appear associated with educational processes in ways that translated into student outcomes on a national licensing examination. Whether informal, internal, interim reviews can mediate this effect remains to be seen.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".