Distinguishing Clinical From Statistical Significances in Contemporary Comparative Effectiveness Research
Bibliographic record
Abstract
OBJECTIVE: To determine the prevalence of clinical significance reporting in contemporary comparative effectiveness research (CER). BACKGROUND: In CER, a statistically significant difference between study groups may or may not be clinically significant. Misinterpreting statistically significant results could lead to inappropriate recommendations that increase health care costs and treatment toxicity. METHODS: CER studies from 2022 issues of the Annals of Surgery , Journal of the American Medical Association , Journal of Clinical Oncology , Journal of Surgical Research , and Journal of the American College of Surgeons were systematically reviewed by 2 different investigators. The primary outcome of interest was whether the authors specified what they considered to be a clinically significant difference in the "Methods." RESULTS: Of 307 reviewed studies, 162 were clinical trials and 145 were observational studies. Authors specified what they considered to be a clinically significant difference in 26 studies (8.5%). Clinical significance was defined using clinically validated standards in 25 studies and subjectively in 1 study. Seven studies (2.3%) recommended a change in clinical decision-making, all with primary outcomes achieving statistical significance. Five (71.4%) of these studies did not have clinical significance defined in their methods. In randomized controlled trials with statistically significant results, sample size was inversely correlated with effect size ( r = -0.30, P = 0.038). CONCLUSIONS: In contemporary CER, most authors do not specify what they consider to be a clinically significant difference in study outcome. Most studies recommending a change in clinical decision-making did so based on statistical significance alone, and clinical significance was usually defined with clinically validated standards.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.088 | 0.018 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".