MétaCan
Menu
Back to cohort
Record W4392104875 · doi:10.1097/sla.0000000000006250

Distinguishing Clinical From Statistical Significances in Contemporary Comparative Effectiveness Research

2024· article· en· W4392104875 on OpenAlexaff
Ajami Gikandi, Julie Hallet, Bas Groot Koerkamp, Clancy J. Clark, Keith D. Lillemoe, Raja R. Narayan, Harvey J. Mamon, Marco A. Zenati, Nabil Wasif, Dana Gelb Safran, Marc G. Besselink, David C. Chang, Lara Traeger, Joel S. Weissman, Zhi Ven Fong

Bibliographic record

VenueAnnals of Surgery · 2024
Typearticle
Languageen
FieldEconomics, Econometrics and Finance
TopicHealth Systems, Economic Evaluations, Quality of Life
Canadian institutionsUniversity of Toronto
FundersNational Cancer Institute
KeywordsMedicineClinical significanceStatistical significanceObservational studyClinical trialSample size determinationComparative effectiveness researchMEDLINEClinical study designInternal medicineIntensive care medicineAlternative medicinePathology

Abstract

fetched live from OpenAlex

OBJECTIVE: To determine the prevalence of clinical significance reporting in contemporary comparative effectiveness research (CER). BACKGROUND: In CER, a statistically significant difference between study groups may or may not be clinically significant. Misinterpreting statistically significant results could lead to inappropriate recommendations that increase health care costs and treatment toxicity. METHODS: CER studies from 2022 issues of the Annals of Surgery , Journal of the American Medical Association , Journal of Clinical Oncology , Journal of Surgical Research , and Journal of the American College of Surgeons were systematically reviewed by 2 different investigators. The primary outcome of interest was whether the authors specified what they considered to be a clinically significant difference in the "Methods." RESULTS: Of 307 reviewed studies, 162 were clinical trials and 145 were observational studies. Authors specified what they considered to be a clinically significant difference in 26 studies (8.5%). Clinical significance was defined using clinically validated standards in 25 studies and subjectively in 1 study. Seven studies (2.3%) recommended a change in clinical decision-making, all with primary outcomes achieving statistical significance. Five (71.4%) of these studies did not have clinical significance defined in their methods. In randomized controlled trials with statistically significant results, sample size was inversely correlated with effect size ( r = -0.30, P = 0.038). CONCLUSIONS: In contemporary CER, most authors do not specify what they consider to be a clinically significant difference in study outcome. Most studies recommending a change in clinical decision-making did so based on statistical significance alone, and clinical significance was usually defined with clinically validated standards.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.088
metaresearch head score (Gemma)0.018
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.096
Threshold uncertainty score0.990

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0880.018
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0020.000
Bibliometrics0.0010.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.970
GPT teacher head0.676
Teacher spread0.294 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations22
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueAnnals of SurgerySame topicHealth Systems, Economic Evaluations, Quality of LifeFrench-language works237,207