MétaCan
Menu
Back to cohort
Record W4392104875 · doi:10.1097/sla.0000000000006250

Distinguishing Clinical From Statistical Significances in Contemporary Comparative Effectiveness Research

2024· article· en· W4392104875 on OpenAlexaff
Ajami Gikandi, Julie Hallet, Bas Groot Koerkamp, Clancy J. Clark, Keith D. Lillemoe, Raja R. Narayan, Harvey J. Mamon, Marco A. Zenati, Nabil Wasif, Dana Gelb Safran, Marc G. Besselink, David C. Chang, Lara Traeger, Joel S. Weissman, Zhi Ven Fong

Bibliographic record

VenueAnnals of Surgery · 2024
Typearticle
Languageen
FieldEconomics, Econometrics and Finance
TopicHealth Systems, Economic Evaluations, Quality of Life
Canadian institutionsUniversity of Toronto
FundersNational Cancer Institute
KeywordsMedicineClinical significanceStatistical significanceObservational studyClinical trialSample size determinationComparative effectiveness researchMEDLINEClinical study designInternal medicineIntensive care medicineAlternative medicinePathology

Abstract

fetched live from OpenAlex

OBJECTIVE: To determine the prevalence of clinical significance reporting in contemporary comparative effectiveness research (CER). BACKGROUND: In CER, a statistically significant difference between study groups may or may not be clinically significant. Misinterpreting statistically significant results could lead to inappropriate recommendations that increase health care costs and treatment toxicity. METHODS: CER studies from 2022 issues of the Annals of Surgery , Journal of the American Medical Association , Journal of Clinical Oncology , Journal of Surgical Research , and Journal of the American College of Surgeons were systematically reviewed by 2 different investigators. The primary outcome of interest was whether the authors specified what they considered to be a clinically significant difference in the "Methods." RESULTS: Of 307 reviewed studies, 162 were clinical trials and 145 were observational studies. Authors specified what they considered to be a clinically significant difference in 26 studies (8.5%). Clinical significance was defined using clinically validated standards in 25 studies and subjectively in 1 study. Seven studies (2.3%) recommended a change in clinical decision-making, all with primary outcomes achieving statistical significance. Five (71.4%) of these studies did not have clinical significance defined in their methods. In randomized controlled trials with statistically significant results, sample size was inversely correlated with effect size ( r = -0.30, P = 0.038). CONCLUSIONS: In contemporary CER, most authors do not specify what they consider to be a clinically significant difference in study outcome. Most studies recommending a change in clinical decision-making did so based on statistical significance alone, and clinical significance was usually defined with clinically validated standards.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.845
metaresearch head score (Gemma)0.947
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.155
Threshold uncertainty score0.191

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.8450.947
Meta-epidemiology (narrow)0.0030.003
Meta-epidemiology (broad)0.0080.009
Bibliometrics0.0300.027
Science and technology studies0.0030.030
Scholarly communication0.0180.020
Open science0.0080.011
Research integrity0.0090.007
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.970
GPT teacher head0.676
Teacher spread0.294 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designTheoretical or conceptual
DomainMethods
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations22
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueAnnals of SurgerySame topicHealth Systems, Economic Evaluations, Quality of LifeFrench-language works237,207