Measuring outcomes in quality improvement education: success is in the eye of the beholder
Bibliographic record
Abstract
Over the past decade, quality improvement (QI) has gone from a secret skill expected only among trained staff in the quality office to a core competency for all health professionals.1–3 This expectation has generated new curricula which have introduced QI to a new generation of learners, but has also created some challenges for health professions educators.4–7 Identifying knowledgeable teachers, defining core content and securing time in the curriculum represent recurring issues, while emerging discussions now centre on how best to evaluate educational efforts in QI. It is here that we find ourselves at an impasse. In this issue of BMJ Quality and Safety, O’Leary and colleagues present their 5-year experience delivering an institutionally sponsored, team-based QI training programme which included attending physicians, residents and fellows and frontline interprofessional team members. They report on its impact on both learner outcomes and project outcomes.8 Their programme demonstrated improvements in participant knowledge, with 172 individuals comprising 32 teams reporting that they had applied their new knowledge and skills to improve clinical quality (87%) and implement QI interventions (62%) at 6 months. At 18 months, nearly half reported leading other QI projects (48%) and many had provided QI mentorship to others (41%). In addition to measuring these learner-focused outcomes, the authors summarise QI project outcomes at programme completion, 6 and 18 months. At one or more of these time points, 20 out of 32 projects (63%) had positive results, defined as showing improvement in one or more project measures without any measure declining in performance. This comprehensive programme evaluation, which includes both learner and project outcomes, provides a unique opportunity to reflect on the goals of QI education for the field of health professions education. Before reflecting on the goals of QI education specifically, it is important to review the yardstick by which best …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".