Notice bibliographique
Résumé
“Take nothing on its looks; take everything on evidence. There’s no better rule.” ─ Charles Dickens, Great Expectations In this issue of the journal, The Open Mind article authored by Drs. Prielipp and Coursin1 provides 3 examples in perioperative medicine wherein initial high-profile research studies have been used to justify changes in public policy, yet, over time, these studies could not be replicated. It is speculated that the evidence-based changes in policy or performance measures, originally based on these studies, have led to deleterious effects, which these authors colorfully compare with the rise in organized crime in response to prohibition. This is a well-written, entertaining, and thought-provoking history of our specialty. We should heed their message and at the same time be aware that although we have great expectations for clinical investigation, and evidence-based medicine, by its nature, it has numerous important limitations. High-profile studies are used to formulate the guideline-directed medical therapy (and public policy). However, high profile does not necessarily equate to high quality. There are numerous examples of studies other than those discussed in this Open Mind article.2–5 Although the need to generate practice guidelines and policy is and rooted in the will to improve patient safety, the quality of the evidence on which many of our current guidelines are formulated may be overrated. In a somewhat controversial article, Kavanagh6 pointed out that the practice of grading evidence was, paradoxically, not evidence-based. Most readers, and reviewers for that matter, of perioperative research rely too heavily of the size of the P value to judge the veracity of a research result,7 whereas credibility should, instead, rely on the design, the effect size, and the biases inherent to the trial. At the American Society of Anesthesiologists’ Annual Meeting in San Francisco in 2013, John Ioannidis gave a guest lecture on medical evidence to the editorial board members of both journals Anesthesia & Analgesia and Anesthesiology. Dr. Ioannidis is an epidemiologist and meta-analyst, whose publications have made him a world-renowned and sought-after lecturer on the credibility of medical evidence. In his article entitled “Why Most Published Research Findings Are False” published in PLoS,8 which, parenthetically, is the most highly cited article in that journal’s history, Dr. Ioannidis has showed, using mathematical modeling, that the probability that the results of a pristinely conducted research study being true is related to 3 factors: the plausibility of the hypothesis being tested (here he uses the term pretest probability); the statistical power of the study (which he points out is related to 2 factors: the samples size and the effect size); and the level of statistical significance. His models suggest that if a study is being conducted on a biologically plausible premise, has a large effect size, and a large sample size (think thousands), then the results of such a study have only an 85% probability of being confirmed as being a true result. This is a classic exercise in logic: by assuming all factors in favor of a theory, the results reveal any inherent weaknesses of the postulate. The article goes on to enumerate several factors, common to medical science, which will decrease the probability of a research finding being true. He notes that the probability of finding a statistically significant result increases by random chance as the number of different teams working on similar problems increases. This may be problematic because it is more likely that positive studies will be published than negative studies.9 In addition, Ioannidis points out that the outcome measured is critically important; although death is an unequivocal event, rating scales and composite outcomes can be manipulated to cast interventions in a better light.10 Surrogate outcomes are common, and we forget that deep vein thrombosis and cardiac ischemia are surrogates for harder, less frequent outcomes. Notably, for harder outcomes, the effect size of an intervention is usually smaller. Ioannidis was also careful to show that his idealized equations used for modeling did not account for the possibility of bias. Bias is always present, at least subtly, in even the best quality blinded randomized clinical trial. In a related article, his team was able to identify >200 types of bias germane to medical science.11 Bias reduces the likelihood that study results are in fact true. Examples of these biases include, but are not limited to: the control group may be have an inflated event rate—e.g., Dutch Echocardiographic Cardiac Risk Evaluation Applying Stress Echocardiography (DECREASE 1)12; failure of randomization to create equitable cohorts—e.g., TRICC13; academic conflicts of interest or outright fraud—e.g., Dr. Boldt,14 Dr. Reuben,a and Dr. Fujii15; selective reporting of outcomes and/or suppression of competing evidence via the peer review process: for instance, Metoprolol after Vascular Surgery (MaVS)16 and POISE117 were both rejected by The New England Journal of Medicine despite the publication of the 2 original high-profile β-blocker studies; and financial incentives.18 The theory espoused in “Why Most Published Research Findings Are False” was then followed by a study entitled “Contradicted and Initially Stronger Effects in Highly Cited Clinical Research,” a real-life confirmation of his previously presented mathematical framework.19 In this study, a search of major medical journals found 49 articles with >1000 citations. A secondary search was then performed for each of the 49 studies to identify subsequent studies on the same subject. Four of the 49 studies were eliminated because they showed no efficacy and also refuted claims of efficacy based on observational studies. In the remaining 45 studies, 14 (33%) either contradicted the initial efficacy claim or found an efficacy of reduced effect size. These subsequent and contradictory studies were found to have both superior design and larger sample sizes, where the median sample size was >3 times larger (2165 subjects). At this point, one should note that the great majority of perioperative medicine research studies and those used in this Open Mind article fall well short of these standards. Ioannidis’ evidence of evidence clearly shows that claims of efficacy (and hence public policy) should be based solely on replicated, minimally biased studies. It also gives us clear direction on how this goal will be achieved; decision-making, high-quality evidence can only come from large, well-designed, prospective trials pursuing plausible theses. Unfortunately, and restated for emphasis, there is a paucity of these examples in perioperative medicine. It is our hope that Drs. Prielipp and Coursin are documenting a process where our specialty is evolving into a high-quality, evidence-based practice. The myriad of pressing perioperative issues such as obstructive sleep apnea, perioperative fluid therapy, appropriate transfusion policies, postoperative renal failure, and the toxicity of anesthetics at extremes of life, to name a few, requires quality evidence for decision-making.8 This evidence can only be attained through the arduous, time-consuming, expensive, and frequently unsuccessful process of clinical investigation. These are indeed great expectations for clinical investigation, and our pursuit of patient safety is a clear justification for its expanded presence. It will take seismic shifts in our behavior to overcome the numerous obstacles that stand in the way of this progress. We must overcome the paucity of research networks, adopt new models for recruiting high volumes of patients, overcome funding issues with new entrepreneurial designs, and shift our focus away from individual academic achievement to reward collaboration. Although many readers may be concerned by the vignettes portrayed in The Open Mind as representing steps in the wrong direction, to my mind they actually highlight the great strength of the evidence-based medicine process: the ability to improve with change.20 DISCLOSURES Name: W. Scott Beattie, MD, PhD, FRCPC. Contribution: This author wrote the manuscript. Attestation: W. Scott Beattie approved the final manuscript. This manuscript was handled by: Sorin J. Brull, MD, FCARCSI (Hon.).
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,022 | 0,160 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,004 | 0,008 |
| Communication savante | 0,018 | 0,013 |
| Science ouverte | 0,004 | 0,010 |
| Intégrité de la recherche | 0,014 | 0,021 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,205 | 0,186 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».