MétaCan
Menu
Back to cohort
Record W2412540422 · doi:10.1213/ane.0000000000000870

Great Expectations

2015· editorial· en· W2412540422 on OpenAlexaff
W. Scott Beattie

Bibliographic record

VenueAnesthesia & Analgesia · 2015
Typeeditorial
Languageen
FieldMedicine
TopicCardiac, Anesthesia and Surgical Outcomes
Canadian institutionsUniversity of TorontoUniversity Health Network
Fundersnot available
KeywordsNothingGrading (engineering)Evidence-based medicineGuidelinePsychologyQuality (philosophy)Public relationsPositive economicsQuality of evidencePolitical scienceEngineering ethicsLaw and economicsMedicineEpistemologyMEDLINELawEconomicsEngineeringPhilosophy

Abstract

fetched live from OpenAlex

“Take nothing on its looks; take everything on evidence. There’s no better rule.” ─ Charles Dickens, Great Expectations In this issue of the journal, The Open Mind article authored by Drs. Prielipp and Coursin1 provides 3 examples in perioperative medicine wherein initial high-profile research studies have been used to justify changes in public policy, yet, over time, these studies could not be replicated. It is speculated that the evidence-based changes in policy or performance measures, originally based on these studies, have led to deleterious effects, which these authors colorfully compare with the rise in organized crime in response to prohibition. This is a well-written, entertaining, and thought-provoking history of our specialty. We should heed their message and at the same time be aware that although we have great expectations for clinical investigation, and evidence-based medicine, by its nature, it has numerous important limitations. High-profile studies are used to formulate the guideline-directed medical therapy (and public policy). However, high profile does not necessarily equate to high quality. There are numerous examples of studies other than those discussed in this Open Mind article.2–5 Although the need to generate practice guidelines and policy is and rooted in the will to improve patient safety, the quality of the evidence on which many of our current guidelines are formulated may be overrated. In a somewhat controversial article, Kavanagh6 pointed out that the practice of grading evidence was, paradoxically, not evidence-based. Most readers, and reviewers for that matter, of perioperative research rely too heavily of the size of the P value to judge the veracity of a research result,7 whereas credibility should, instead, rely on the design, the effect size, and the biases inherent to the trial. At the American Society of Anesthesiologists’ Annual Meeting in San Francisco in 2013, John Ioannidis gave a guest lecture on medical evidence to the editorial board members of both journals Anesthesia & Analgesia and Anesthesiology. Dr. Ioannidis is an epidemiologist and meta-analyst, whose publications have made him a world-renowned and sought-after lecturer on the credibility of medical evidence. In his article entitled “Why Most Published Research Findings Are False” published in PLoS,8 which, parenthetically, is the most highly cited article in that journal’s history, Dr. Ioannidis has showed, using mathematical modeling, that the probability that the results of a pristinely conducted research study being true is related to 3 factors: the plausibility of the hypothesis being tested (here he uses the term pretest probability); the statistical power of the study (which he points out is related to 2 factors: the samples size and the effect size); and the level of statistical significance. His models suggest that if a study is being conducted on a biologically plausible premise, has a large effect size, and a large sample size (think thousands), then the results of such a study have only an 85% probability of being confirmed as being a true result. This is a classic exercise in logic: by assuming all factors in favor of a theory, the results reveal any inherent weaknesses of the postulate. The article goes on to enumerate several factors, common to medical science, which will decrease the probability of a research finding being true. He notes that the probability of finding a statistically significant result increases by random chance as the number of different teams working on similar problems increases. This may be problematic because it is more likely that positive studies will be published than negative studies.9 In addition, Ioannidis points out that the outcome measured is critically important; although death is an unequivocal event, rating scales and composite outcomes can be manipulated to cast interventions in a better light.10 Surrogate outcomes are common, and we forget that deep vein thrombosis and cardiac ischemia are surrogates for harder, less frequent outcomes. Notably, for harder outcomes, the effect size of an intervention is usually smaller. Ioannidis was also careful to show that his idealized equations used for modeling did not account for the possibility of bias. Bias is always present, at least subtly, in even the best quality blinded randomized clinical trial. In a related article, his team was able to identify >200 types of bias germane to medical science.11 Bias reduces the likelihood that study results are in fact true. Examples of these biases include, but are not limited to: the control group may be have an inflated event rate—e.g., Dutch Echocardiographic Cardiac Risk Evaluation Applying Stress Echocardiography (DECREASE 1)12; failure of randomization to create equitable cohorts—e.g., TRICC13; academic conflicts of interest or outright fraud—e.g., Dr. Boldt,14 Dr. Reuben,a and Dr. Fujii15; selective reporting of outcomes and/or suppression of competing evidence via the peer review process: for instance, Metoprolol after Vascular Surgery (MaVS)16 and POISE117 were both rejected by The New England Journal of Medicine despite the publication of the 2 original high-profile β-blocker studies; and financial incentives.18 The theory espoused in “Why Most Published Research Findings Are False” was then followed by a study entitled “Contradicted and Initially Stronger Effects in Highly Cited Clinical Research,” a real-life confirmation of his previously presented mathematical framework.19 In this study, a search of major medical journals found 49 articles with >1000 citations. A secondary search was then performed for each of the 49 studies to identify subsequent studies on the same subject. Four of the 49 studies were eliminated because they showed no efficacy and also refuted claims of efficacy based on observational studies. In the remaining 45 studies, 14 (33%) either contradicted the initial efficacy claim or found an efficacy of reduced effect size. These subsequent and contradictory studies were found to have both superior design and larger sample sizes, where the median sample size was >3 times larger (2165 subjects). At this point, one should note that the great majority of perioperative medicine research studies and those used in this Open Mind article fall well short of these standards. Ioannidis’ evidence of evidence clearly shows that claims of efficacy (and hence public policy) should be based solely on replicated, minimally biased studies. It also gives us clear direction on how this goal will be achieved; decision-making, high-quality evidence can only come from large, well-designed, prospective trials pursuing plausible theses. Unfortunately, and restated for emphasis, there is a paucity of these examples in perioperative medicine. It is our hope that Drs. Prielipp and Coursin are documenting a process where our specialty is evolving into a high-quality, evidence-based practice. The myriad of pressing perioperative issues such as obstructive sleep apnea, perioperative fluid therapy, appropriate transfusion policies, postoperative renal failure, and the toxicity of anesthetics at extremes of life, to name a few, requires quality evidence for decision-making.8 This evidence can only be attained through the arduous, time-consuming, expensive, and frequently unsuccessful process of clinical investigation. These are indeed great expectations for clinical investigation, and our pursuit of patient safety is a clear justification for its expanded presence. It will take seismic shifts in our behavior to overcome the numerous obstacles that stand in the way of this progress. We must overcome the paucity of research networks, adopt new models for recruiting high volumes of patients, overcome funding issues with new entrepreneurial designs, and shift our focus away from individual academic achievement to reward collaboration. Although many readers may be concerned by the vignettes portrayed in The Open Mind as representing steps in the wrong direction, to my mind they actually highlight the great strength of the evidence-based medicine process: the ability to improve with change.20 DISCLOSURES Name: W. Scott Beattie, MD, PhD, FRCPC. Contribution: This author wrote the manuscript. Attestation: W. Scott Beattie approved the final manuscript. This manuscript was handled by: Sorin J. Brull, MD, FCARCSI (Hon.).

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.022
metaresearch head score (Gemma)0.160
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: none
Teacher disagreement score0.205
Threshold uncertainty score0.687

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0220.160
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0020.001
Science and technology studies0.0040.008
Scholarly communication0.0180.013
Open science0.0040.010
Research integrity0.0140.021
Insufficient payload (model declined to judge)0.2050.186

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.015
GPT teacher head0.290
Teacher spread0.275 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueAnesthesia & AnalgesiaSame topicCardiac, Anesthesia and Surgical OutcomesFrench-language works237,207