Reasons to Believe: Biostatistics and Methodology for the Neurosurgeon
Notice bibliographique
Résumé
“If I listen long enough to you, I’d find a reason to believe that it's all true.” Tim Hardin, Reason to Believe, 1971 Our human need to believe that we understand the world around us is a powerful motivator. It can also be the source of misconception, superstition, delusion, fallacy, and other forms of misbelief. The development of the fundamental principles of science and, specifically, the methods of objective study design and analysis is essentially the history of understanding the many ways we can fool ourselves into believing that our preconceptions are true. Those principles and methods have similarly been misunderstood and resisted. The underlying commitment to objectivity and truth has sometimes been distorted and the methods used inappropriately to support untruthful ideas. One of the best known expressions of this is the quote: “There are three kinds of lies: lies, damned lies, and statistics.” Variously attributed to Benjamin Disraeli (by Mark Twain), Leonard H. (Lord) Courtney, and Arthur James (Earl) Balfour,1 this oft repeated aphorism has poisoned the minds of many physicians to the study of statistics. If the methods of statistics are a way to lie rather than to reveal truth, why would a person with honorable intention spend time to learn a method of deception? The answer is straightforward – statistical methods are not meant for lying, but it is easy to make inappropriate use of them. Thus, to make proper use of the results of statistical analyses, the reader must be fluent enough with statistical concepts and techniques to identify inappropriate methods and conclusions. Clinical researchers rarely deliberately lie. However, statistical concepts and methods applied to human disease are arguably among the most complicated and difficult to understand both for the physician trying to understand the statistics and the statistician trying to understand the medicine. Misunderstanding and misapplication of these techniques can lead to inadvertent “lies.” Therefore, experts in biostatistical methodology have long sought to educate physicians in the value of their methods. We join that effort today. In this series we hope to demonstrate that, despite the many ways in which biostatistical methods can be misunderstood and misapplied, thoughtful and rigorous design and analysis can provide the reader with many reasons to believe in the published results that inform their daily practice. One of the earliest attempts to statistically educate physicians was a series of articles by Austin Bradford Hill published in the Lancet in 1937.2 More recently, between 1967 and 1969 Donald Mainland,3 Professor and Chair of Medical Statistics at New York University, published a series of articles in Clinical Pharmacology and Therapeutics entitled “Statistical ward rounds.” These were collected in book form.4 They grew out of a series of “Notes From a Laboratory of Medical Statistics” that he distributed to research workers with whom he worked. In 1969 this series was taken over by Alvan Feinstein,5 Sterling Professor of Medicine and Epidemiology at Yale, who published a 57 article series retitled “Clinical Biostatistics” ending in 1981. These delightfully titled articles (for example, “The haze of Bayes, the aerial palaces of decision analysis, and the computerized Ouija board”) were my first formal introduction to what we now call Clinical Epidemiology (the concepts, although not the original articles, were published as a book6). In the 1980s David Sackett and his colleagues at McMaster University published a number of smaller article series on focused topics (for example, “Interpretation of Diagnostic Data” in the Canadian Medical Journal, 1983; “How To Keep Up With the Medical Literature”, Annals of Internal Medicine, 1986). Their work led to the comprehensive series “User's Guide's to the Medical Literature” published in JAMA between 1993 and 2000 with sporadic articles continuing to the present. They have been collected in book form.7 These articles are not focused entirely upon statistics, but rather introduce the reader to the basic and applied concepts of various types of study design required to answer clinical questions. In addition, and in parallel with the effort in JAMA, the British Medical Journal has published “Statistics Notes” from 1994 to the present (for a summary, see the website reference).8 These efforts continue today. There has been a tendency in recent years to concentrate on a particular subset of issues. For example, the current series in the New England Journal of Medicine “The changing face of clinical trials” or “The GRADE Guidelines” in the Journal of Clinical Investigation. Sporadic articles on particular topics appear under the heading “Statistics in Medicine” in the New England Journal of Medicine. Medical subspecialty journals (like ours) provide article series with clinical examples familiar to their clinician audiences (for example, see Wolfe and Abramson9 and “Ophthalmic statistics note”, British Journal of Ophthalmology, ongoing). The need for such educational efforts seems never to be satisfied. In part this is because the methods of designing, conducting, analyzing, interpreting, and summarizing clinical research are continually evolving and therefore becoming more complex. It is also because the topics are devilishly difficult to write about in an interesting and understandable way. The idea that physicians might improve patient care by becoming statistically literate is a relatively new one. The first serious efforts at tabulating results and using those tabulations to inform treatment are ascribed to Pierre Charles Alexandre Louis who practiced in France in the middle of the 19th century and argued his case in the French Academy of Sciences.10 Acceptance was – and continues to be – slow. Many of the classic statistical tests still used today (t-tests, ANOVA) were developed early in the 20th century, a little before Bovie and Cushing introduced electrocautery to the operating room. The first formal randomized trials were done in the middle of the 20th century when most craniotomies were done with a Gigli saw. The transition from observational to experimental designs (originally developed for plants and animals but adapted to clinical trials designed for humans) is a late 20th century development. The immense increase in computing power in the late 20th century now allows widespread access to complex analytic techniques such as cluster, factor, and survival analysis. Multiple logistic regression can be applied to massive databases from governmental and private insurance providers. These increasingly sophisticated ways of extracting signal from noise and meaning from babble also increase the risk that misapplication and faulty interpretation can lead to incorrect conclusions or inadvertent lies. We are neurosurgeons and therefore we intend to master the concepts of excellent clinical research and we are undaunted by the obstacles of time and dense, unamusing content. Therefore, we will undertake a series of articles, intended to be informational and instructive, that highlight the concepts that underlie high-quality clinical research with neurosurgical examples. The goal is not to turn neurosurgeons into biostatisticians, but into sophisticated consumers of biostatistical expertise. Indeed, neurosurgeons should be leery of becoming their own biostatisticians. The field is arcanely complex and it is hard to stay current with the newest techniques. But the neurosurgeon in private or academic practice needs to understand the concepts of study design, analysis, and interpretation well enough to be able to correctly interpret the results presented, even if the authors have misinterpreted their own data. We should be able to ask good questions of biostatistical consultants and identify concerns about methods used in presentations and published papers. In essence, the reader should be able to separate appropriate from inappropriate questions, techniques, and conclusions. In an ideal world we would present an orderly compendium of topics, elegantly written and laced with humor. The Table of Contents would look something like this: Asking Good Clinical Research Questions Collecting Good Evidence Properly Analyzing Evidence Accurately and Responsibly Reporting Evidence Scientifically Summarizing Evidence Appropriately Interpreting Evidence Challenges for the Future This is the real world and if we expect real neurosurgeons to participate in writing these articles we cannot enforce such a schedule. Therefore, we will publish articles that address these topics as they become available. Some will be invited, some may be chosen from articles submitted to the journal through the usual channels. In the end we expect to have a reasonably complete review of these important topics. We have asked neurosurgeons familiar with various aspects of biostatistics and clinical epidemiology to collaborate with biostatisticians who have some experience with clinical neurosurgical research in finding, editing, and writing these essays. Some arise from perceived need (a series of articles on meta-analysis, a powerful and oft abused method of summarizing evidence), some from the list of identified topics (crafting good questions for clinical research), and some may come from studies submitted to the journal through the usual publication process that seem especially appropriate to the goals of this series. We encourage the reader, whether established expert or neurosurgical trainee, to spend the time necessary to read and understand these articles. They offer one of the best ways to protect your patients from poorly conducted or incorrectly interpreted clinical research findings. They also offer a way to manage the ever-increasing burden of published clinical research that the neurosurgeon must review. An article with a poorly designed question or inappropriate research method need not be studied further. We welcome your suggestions, comments, criticisms, and even manuscript submissions. In the nearly 180 years since Louis presented his controversial thesis that “…I conceive that without the aid of statistics nothing like real medical science is possible” no one has found a “best” way to convey the essence of numerical analysis of complex biological data to practicing physicians. We join the effort because it is fundamentally important to providing excellent care to our patients. Disclosure The author has no personal, financial, or institutional interest in any of the drugs, materials, or devices described in this article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,318 | 0,682 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,004 | 0,002 |
| Bibliométrie | 0,011 | 0,011 |
| Études des sciences et des technologies | 0,003 | 0,028 |
| Communication savante | 0,013 | 0,015 |
| Science ouverte | 0,004 | 0,009 |
| Intégrité de la recherche | 0,014 | 0,030 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,004 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».