Reasons to Believe: Biostatistics and Methodology for the Neurosurgeon
Bibliographic record
Abstract
“If I listen long enough to you, I’d find a reason to believe that it's all true.” Tim Hardin, Reason to Believe, 1971 Our human need to believe that we understand the world around us is a powerful motivator. It can also be the source of misconception, superstition, delusion, fallacy, and other forms of misbelief. The development of the fundamental principles of science and, specifically, the methods of objective study design and analysis is essentially the history of understanding the many ways we can fool ourselves into believing that our preconceptions are true. Those principles and methods have similarly been misunderstood and resisted. The underlying commitment to objectivity and truth has sometimes been distorted and the methods used inappropriately to support untruthful ideas. One of the best known expressions of this is the quote: “There are three kinds of lies: lies, damned lies, and statistics.” Variously attributed to Benjamin Disraeli (by Mark Twain), Leonard H. (Lord) Courtney, and Arthur James (Earl) Balfour,1 this oft repeated aphorism has poisoned the minds of many physicians to the study of statistics. If the methods of statistics are a way to lie rather than to reveal truth, why would a person with honorable intention spend time to learn a method of deception? The answer is straightforward – statistical methods are not meant for lying, but it is easy to make inappropriate use of them. Thus, to make proper use of the results of statistical analyses, the reader must be fluent enough with statistical concepts and techniques to identify inappropriate methods and conclusions. Clinical researchers rarely deliberately lie. However, statistical concepts and methods applied to human disease are arguably among the most complicated and difficult to understand both for the physician trying to understand the statistics and the statistician trying to understand the medicine. Misunderstanding and misapplication of these techniques can lead to inadvertent “lies.” Therefore, experts in biostatistical methodology have long sought to educate physicians in the value of their methods. We join that effort today. In this series we hope to demonstrate that, despite the many ways in which biostatistical methods can be misunderstood and misapplied, thoughtful and rigorous design and analysis can provide the reader with many reasons to believe in the published results that inform their daily practice. One of the earliest attempts to statistically educate physicians was a series of articles by Austin Bradford Hill published in the Lancet in 1937.2 More recently, between 1967 and 1969 Donald Mainland,3 Professor and Chair of Medical Statistics at New York University, published a series of articles in Clinical Pharmacology and Therapeutics entitled “Statistical ward rounds.” These were collected in book form.4 They grew out of a series of “Notes From a Laboratory of Medical Statistics” that he distributed to research workers with whom he worked. In 1969 this series was taken over by Alvan Feinstein,5 Sterling Professor of Medicine and Epidemiology at Yale, who published a 57 article series retitled “Clinical Biostatistics” ending in 1981. These delightfully titled articles (for example, “The haze of Bayes, the aerial palaces of decision analysis, and the computerized Ouija board”) were my first formal introduction to what we now call Clinical Epidemiology (the concepts, although not the original articles, were published as a book6). In the 1980s David Sackett and his colleagues at McMaster University published a number of smaller article series on focused topics (for example, “Interpretation of Diagnostic Data” in the Canadian Medical Journal, 1983; “How To Keep Up With the Medical Literature”, Annals of Internal Medicine, 1986). Their work led to the comprehensive series “User's Guide's to the Medical Literature” published in JAMA between 1993 and 2000 with sporadic articles continuing to the present. They have been collected in book form.7 These articles are not focused entirely upon statistics, but rather introduce the reader to the basic and applied concepts of various types of study design required to answer clinical questions. In addition, and in parallel with the effort in JAMA, the British Medical Journal has published “Statistics Notes” from 1994 to the present (for a summary, see the website reference).8 These efforts continue today. There has been a tendency in recent years to concentrate on a particular subset of issues. For example, the current series in the New England Journal of Medicine “The changing face of clinical trials” or “The GRADE Guidelines” in the Journal of Clinical Investigation. Sporadic articles on particular topics appear under the heading “Statistics in Medicine” in the New England Journal of Medicine. Medical subspecialty journals (like ours) provide article series with clinical examples familiar to their clinician audiences (for example, see Wolfe and Abramson9 and “Ophthalmic statistics note”, British Journal of Ophthalmology, ongoing). The need for such educational efforts seems never to be satisfied. In part this is because the methods of designing, conducting, analyzing, interpreting, and summarizing clinical research are continually evolving and therefore becoming more complex. It is also because the topics are devilishly difficult to write about in an interesting and understandable way. The idea that physicians might improve patient care by becoming statistically literate is a relatively new one. The first serious efforts at tabulating results and using those tabulations to inform treatment are ascribed to Pierre Charles Alexandre Louis who practiced in France in the middle of the 19th century and argued his case in the French Academy of Sciences.10 Acceptance was – and continues to be – slow. Many of the classic statistical tests still used today (t-tests, ANOVA) were developed early in the 20th century, a little before Bovie and Cushing introduced electrocautery to the operating room. The first formal randomized trials were done in the middle of the 20th century when most craniotomies were done with a Gigli saw. The transition from observational to experimental designs (originally developed for plants and animals but adapted to clinical trials designed for humans) is a late 20th century development. The immense increase in computing power in the late 20th century now allows widespread access to complex analytic techniques such as cluster, factor, and survival analysis. Multiple logistic regression can be applied to massive databases from governmental and private insurance providers. These increasingly sophisticated ways of extracting signal from noise and meaning from babble also increase the risk that misapplication and faulty interpretation can lead to incorrect conclusions or inadvertent lies. We are neurosurgeons and therefore we intend to master the concepts of excellent clinical research and we are undaunted by the obstacles of time and dense, unamusing content. Therefore, we will undertake a series of articles, intended to be informational and instructive, that highlight the concepts that underlie high-quality clinical research with neurosurgical examples. The goal is not to turn neurosurgeons into biostatisticians, but into sophisticated consumers of biostatistical expertise. Indeed, neurosurgeons should be leery of becoming their own biostatisticians. The field is arcanely complex and it is hard to stay current with the newest techniques. But the neurosurgeon in private or academic practice needs to understand the concepts of study design, analysis, and interpretation well enough to be able to correctly interpret the results presented, even if the authors have misinterpreted their own data. We should be able to ask good questions of biostatistical consultants and identify concerns about methods used in presentations and published papers. In essence, the reader should be able to separate appropriate from inappropriate questions, techniques, and conclusions. In an ideal world we would present an orderly compendium of topics, elegantly written and laced with humor. The Table of Contents would look something like this: Asking Good Clinical Research Questions Collecting Good Evidence Properly Analyzing Evidence Accurately and Responsibly Reporting Evidence Scientifically Summarizing Evidence Appropriately Interpreting Evidence Challenges for the Future This is the real world and if we expect real neurosurgeons to participate in writing these articles we cannot enforce such a schedule. Therefore, we will publish articles that address these topics as they become available. Some will be invited, some may be chosen from articles submitted to the journal through the usual channels. In the end we expect to have a reasonably complete review of these important topics. We have asked neurosurgeons familiar with various aspects of biostatistics and clinical epidemiology to collaborate with biostatisticians who have some experience with clinical neurosurgical research in finding, editing, and writing these essays. Some arise from perceived need (a series of articles on meta-analysis, a powerful and oft abused method of summarizing evidence), some from the list of identified topics (crafting good questions for clinical research), and some may come from studies submitted to the journal through the usual publication process that seem especially appropriate to the goals of this series. We encourage the reader, whether established expert or neurosurgical trainee, to spend the time necessary to read and understand these articles. They offer one of the best ways to protect your patients from poorly conducted or incorrectly interpreted clinical research findings. They also offer a way to manage the ever-increasing burden of published clinical research that the neurosurgeon must review. An article with a poorly designed question or inappropriate research method need not be studied further. We welcome your suggestions, comments, criticisms, and even manuscript submissions. In the nearly 180 years since Louis presented his controversial thesis that “…I conceive that without the aid of statistics nothing like real medical science is possible” no one has found a “best” way to convey the essence of numerical analysis of complex biological data to practicing physicians. We join the effort because it is fundamentally important to providing excellent care to our patients. Disclosure The author has no personal, financial, or institutional interest in any of the drugs, materials, or devices described in this article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.318 | 0.682 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.002 |
| Bibliometrics | 0.011 | 0.011 |
| Science and technology studies | 0.003 | 0.028 |
| Scholarly communication | 0.013 | 0.015 |
| Open science | 0.004 | 0.009 |
| Research integrity | 0.014 | 0.030 |
| Insufficient payload (model declined to judge) | 0.006 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".