Evaluation of Continuing Professional Development for Physicians – Time for Change: A Scoping Review
Bibliographic record
Abstract
Introduction: Evaluation of education interventions is essential for continuous improvement as it provides insights into how and why outcomes occur. Specifically, for physicians' continuing professional development (CPD) programs, which aim to upskill physicians in a range of practice-essential domains, evaluations are crucial to assure physicians' continuous development, enhanced patient care and safety. However, evaluations of health professions education (HPE) interventions tend to be outcomes focused, failing to capture how and why outcomes occur. This scoping review aimed to identify evaluation techniques used to evaluate CPD programs for physicians, and to determine how these techniques are being implemented as well as the their quality. Methods: We searched PubMed, Embase, Web of Science, among others for English publications on evaluation of CPD programs for physicians, in the past decade. We used a data charting template to extract study details regarding the evaluation techniques and produced a checklist to assess the quality of the evaluations. Results: 101 studies were included; of which 91 studies did not use an evaluation framework. Our findings revealed shortcomings in the evaluations of CPD programs including lack of attention to: intervention processes; unintended outcomes and contextual factors; use of theory; evaluation framework use; and rationale for chosen evaluation method. Discussion: Our findings highlighted major gaps in the evaluation techniques employed in physicians' CPD. Attention needs to be paid to evaluating both program processes and outcomes to illuminate how and why impacts are or are not occurring.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.037 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".