Bibliographic record
Abstract
While most of us were enjoying the pleasures of an idle summer, a publication in Mayo Clinic Proceedings created a stir in the esoteric world of evidence-based medicine. The subject of interest was medical reversal, a phenomenon in which ‘‘a medical practice is found to be inferior to some lesser or prior standard of care’’. Prasad et al. reviewed original articles published in the New England Journal of Medicine from 2001 to 2010 in their endeavour to determine if new evidence advanced, confirmed, or rejected current medical practices. Seven hundred fiftysix (56%) of the 1,344 manuscripts reviewed found a new therapy superior to an older treatment, 138 (10%) confirmed the utility of a current therapy, and 165 (12%) found new practices inferior to present care. One hundred forty-six (40%) of the 363 studies evaluating an existing medical practice found the current therapy inferior to a lesser or earlier standard of care. These rejections of current practice for an older treatment (or no treatment at all) cut across classes of medical care, including anesthesia. Cited examples from perioperative medicine included bispectral index monitoring, mild hypothermia for an intracranial aneurysm clipping, use of a pulmonary artery catheter for high-risk surgical patients, coronary revascularization before elective vascular surgery, epidurals in early labour, and use of aprotinin in cardiac surgery. The reaction to Prasad’s findings was swift and critical. An accompanying editorial stated, ‘‘[T]he proportion of medical reversals seems alarmingly high. At a minimum, it poses major questions about the validity and clinical utility of a sizeable portion of everyday medical care.’’ The blog, Science-Based Medicine, noted, ‘‘[T]his highlights the fact that some current practices are useless or less than optimal and need to be reexamined.’’ Even the New York Times got in on the action by leading with the headline, ‘‘Medical Procedures May Be Useless, or Worse.’’ Both the evidence base for medical practice and the practice itself were under attack. Patients and clinicians alike could be left wondering how a supposedly science-based practice could have been wrong so frequently. Prasad identified a common narrative among the reversals noting that, ‘‘[a]lthough there is a weak evidence base for some practice, it gains acceptance largely through vocal support from prominent advocates and faith that the mechanism of action is sound. Later, future trials undermine the therapy, but removing the contradicted practice often proves challenging.’’ Does this sound familiar? Let us consider the saga of perioperative beta-blockade from Prasad’s perspective of medical reversal. In 1996, Mangano published results of a 200-patient placebo-controlled trial evaluating a seven-day perioperative course of atenolol on a composite outcome of cardiac mortality and morbidity. Six patients who died in hospital were excluded. Of those surviving to hospital discharge, 12 of 99 placebo patients died within six months of surgery compared with four of 95 patients receiving atenolol (crude relative risk [RR] 0.35; 95% confidence interval [CI] 0.1 to 1.0). Several years later, Poldermans G. L. Bryson, MD (&) Department of Anesthesiology, The Ottawa Hospital, The University of Ottawa, 1053 Carling Avenue, Box 249C, Ottawa, ON K1Y 4E9, Canada e-mail: glbryson@ottawahospital.on.ca
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.074 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.004 | 0.010 |
| Scholarly communication | 0.012 | 0.013 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.020 | 0.040 |
| Insufficient payload (model declined to judge) | 0.006 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".