Evaluation of continuous quality improvement in accreditation for medical education
Bibliographic record
Abstract
BACKGROUND: Accreditation systems are based on a number of principles and purposes that vary across jurisdictions. Decision making about accreditation governance suffers from a paucity of evidence. This paper evaluates the pros and cons of continuous quality improvement (CQI) within educational institutions that have traditionally been accredited based on episodic evaluation by external reviewers. METHODS: A naturalistic utility-focused evaluation was performed. Seven criteria, each relevant to government oversight, were used to evaluate the pros and cons of the use of CQI in three medical school accreditation systems across the continuum of medical education. The authors, all involved in the governance of accreditation, iteratively discussed CQI in their medical education contexts in light of the seven criteria until consensus was reached about general patterns. RESULTS: Because institutional CQI makes use of early warning systems, it may enhance the reflective function of accreditation. In the three medical accreditation systems examined, external accreditors lacked the ability to respond quickly to local events or societal developments. There is a potential role for CQI in safeguarding the public interest. Moreover, the central governance structure of accreditation may benefit from decentralized CQI. However, CQI has weaknesses with respect to impartiality, independence, and public accountability, as well as with the ability to balance expectations with capacity. CONCLUSION: CQI, as evaluated with the seven criteria of oversight, has pros and cons. Its use still depends on the balance between the expected positive effects-especially increased reflection and faster response to important issues-versus the potential impediments. A toxic culture that affects impartiality and independence, as well as the need to invest in bureaucratic systems may make in impractical for some institutions to undertake CQI.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.139 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".