Rigorous evaluations of evolving interventions: can we have our cake and eat it too?
Bibliographic record
Abstract
The years immediately following the widespread interest in patient safety1 and then healthcare quality2 saw considerable debate between pragmatically oriented improvers and research-oriented evaluators3–6 —or between ‘evangelists’ and ‘snails’ as one longtime observer characterised the two groups.7 Too often, enthusiastic improvers (‘evangelists’) relied on simple pre-post designs within a single context leading to erroneous claims of efficacy.8 In contrast, research-oriented investigators (‘snails’) and journals pushed for ever more rigorous designs including randomised trials, potentially at the cost of discouraging many improvers without this training and leading to slower development and deployment of effective interventions.9 10 Many clinicians, quality improvement (QI) experts and researchers are thus caught in a quandary: how best to evaluate a candidate QI intervention? How can we best balance the pragmatic needs of improvement—including the frequent need to refine the intervention or its implementation—with the requirement of most traditional evaluative designs, which typically require a static intervention? We believe this question is one of the most important issues to consider when developing a QI intervention and is often not considered carefully enough—either by snails or evangelists. Decisions about when and how to evaluate potentially promising interventions can have crucial implications for the future of the intervention and the patients it could affect. In this issue of BMJ Quality and Safety , Swaminathan and colleagues11 present a rigorous evaluation of the Michigan Appropriateness Guide for Intravenous Catheters (MAGIC) QI intervention, intended to reduce adverse events stemming from the insertion of peripherally inserted venous central catheters (PICC). PICCs have become ubiquitous as a substitution for a central intravenous line when patients need longer term central intravenous access, but clinicians often order them unnecessarily or order inappropriate types—for example, a double-lumen PICC when a single-lumen PICC would work just as well and carry a lower risk of complications. …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.620 | 0.832 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.008 | 0.007 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.007 | 0.036 |
| Scholarly communication | 0.026 | 0.035 |
| Open science | 0.012 | 0.014 |
| Research integrity | 0.026 | 0.033 |
| Insufficient payload (model declined to judge) | 0.007 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".