MétaCan
Menu
Back to cohort
Record W2792950363 · doi:10.1136/bmjqs-2017-007554

Rigorous evaluations of evolving interventions: can we have our cake and eat it too?

2018· letter· en· W2792950363 on OpenAlexaff
Robert E. Burke, Kaveh G Shojania

Bibliographic record

VenueBMJ Quality & Safety · 2018
Typeletter
Languageen
FieldEconomics, Econometrics and Finance
TopicHealth Systems, Economic Evaluations, Quality of Life
Canadian institutionsHealth Sciences CentreUniversity of TorontoSunnybrook Health Science Centre
FundersU.S. Department of Veterans Affairs
KeywordsPsychological interventionMedicineContext (archaeology)Intervention (counseling)Health careQuality (philosophy)Best practiceSoftware deploymentMagic bulletNursingComputer scienceManagement

Abstract

fetched live from OpenAlex

The years immediately following the widespread interest in patient safety1 and then healthcare quality2 saw considerable debate between pragmatically oriented improvers and research-oriented evaluators3–6 —or between ‘evangelists’ and ‘snails’ as one longtime observer characterised the two groups.7 Too often, enthusiastic improvers (‘evangelists’) relied on simple pre-post designs within a single context leading to erroneous claims of efficacy.8 In contrast, research-oriented investigators (‘snails’) and journals pushed for ever more rigorous designs including randomised trials, potentially at the cost of discouraging many improvers without this training and leading to slower development and deployment of effective interventions.9 10 Many clinicians, quality improvement (QI) experts and researchers are thus caught in a quandary: how best to evaluate a candidate QI intervention? How can we best balance the pragmatic needs of improvement—including the frequent need to refine the intervention or its implementation—with the requirement of most traditional evaluative designs, which typically require a static intervention? We believe this question is one of the most important issues to consider when developing a QI intervention and is often not considered carefully enough—either by snails or evangelists. Decisions about when and how to evaluate potentially promising interventions can have crucial implications for the future of the intervention and the patients it could affect. In this issue of BMJ Quality and Safety , Swaminathan and colleagues11 present a rigorous evaluation of the Michigan Appropriateness Guide for Intravenous Catheters (MAGIC) QI intervention, intended to reduce adverse events stemming from the insertion of peripherally inserted venous central catheters (PICC). PICCs have become ubiquitous as a substitution for a central intravenous line when patients need longer term central intravenous access, but clinicians often order them unnecessarily or order inappropriate types—for example, a double-lumen PICC when a single-lumen PICC would work just as well and carry a lower risk of complications. …

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.620
metaresearch head score (Gemma)0.832
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.380
Threshold uncertainty score0.469

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.6200.832
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0080.007
Bibliometrics0.0060.006
Science and technology studies0.0070.036
Scholarly communication0.0260.035
Open science0.0120.014
Research integrity0.0260.033
Insufficient payload (model declined to judge)0.0070.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.550
GPT teacher head0.532
Teacher spread0.018 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations19
Published2018
Admission routes1
Has abstractyes

Explore more

Same venueBMJ Quality & SafetySame topicHealth Systems, Economic Evaluations, Quality of LifeFrench-language works237,207