MétaCan
Menu
Back to cohort
Record W4210362394 · doi:10.1111/jep.13657

High quality (certainty) evidence changes less often than low‐quality evidence, but the magnitude of effect size does not systematically differ between studies with low versus high‐quality evidence

2022· article· en· W4210362394 on OpenAlexaff
Benjamin Djulbegović, Muhammad Muneeb Ahmed, Iztok Hozo, Despina Koletsi, Lars G. Hemkens, Amy Price, Rachel Riera, Paulo Nadanovsky, Ana Paula Pires dos Santos, Daniela Oliveira de Melo, Ranjan Pathak, Rafael Leite Pacheco, Luis Eduardo Santos Fontes, Enderson Miranda, David Nunan

Bibliographic record

VenueJournal of Evaluation in Clinical Practice · 2022
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsMcMaster University
FundersAgency for Healthcare Research and Quality
KeywordsMedicineGrading (engineering)Evidence-based medicineStatisticsSystematic reviewPublication biasMeta-analysisStandard deviationEconometricsMEDLINEMathematicsInternal medicineAlternative medicine

Abstract

fetched live from OpenAlex

Abstract Rationale, Aims, and Objectives It is generally believed that evidence from low quality of evidence generate inaccurate estimates about treatment effects more often than evidence from high (certainty) quality evidence (CoE). As a result, we would expect that (a) estimates of effects of health interventions initially based on high CoE change less frequently than the effects estimated by lower CoE (b) the estimates of magnitude of effect size differ between high and low CoE. Empirical assessment of these foundational principles of evidence‐based medicine has been lacking. Methods We reviewed the Cochrane Database of Systematic Reviews from January 2016 through May 2021 for pairs of original and updated reviews for change in CoE assessments based on the Grading of Recommendations Assessment, Development and Evaluation (GRADE) method. We assessed the difference in effect sizes between the original versus updated reviews as a function of change in CoE, which we report as a ratio of odds ratio (ROR). We compared ROR generated in the studies in which CoE changed from very low/low (VL/L) to moderate/high (M/H) versus M/H to VL/L. Heterogeneity and inconsistency were assessed using the tau andI2statistic. We also assessed the change in precision of effect estimates (by calculating the ratio of standard errors) (seR), and the absolute deviation in estimates of treatment effects (aROR). Results Four hundred and nineteen pairs of reviews were included of which 414 (207 × 2) informed the CoE appraisal and 384 (192 × 2) the assessment of effect size. We found that CoE originally appraised as VL/L had 2.1 [95% confidence interval (CI): 1.19–4.12;p = 0.0091] times higher odds to be changed in the future studies than M/H CoE. However, the effect size was not different (p = 1) when CoE changed from VL/L → M/H [ROR = 1.02 (95% CI: 0.74–1.39)] compared with M/H → VL/L (ROR = 1.02 [95% CI: 0.44–2.37]). Similar overlap in aROR between the VL/L → M/H versus M/H → VL/L subgroups was observed [median (IQR): 1.12 (1.07–1.57) vs. 1.21 (1.12–2.43)]. We observed large inconsistency across ROR estimates (I2 = 99%). There was larger imprecision in treatment effects when CoE changed from VL/L → M/H (seR = 1.46) than when it changed from M/H → VL/L (seR = 0.72). Conclusions We found that low‐quality evidence changes more often than high CoE. However, the effect size did not systematically differ between the studies with low versus high CoE. The finding that the effect size did not differ between low and high CoE indicate urgent need to refine current EBM critical appraisal methods.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.287
metaresearch head score (Gemma)0.713
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.713
Threshold uncertainty score0.879

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.2870.713
Meta-epidemiology (narrow)0.0020.004
Meta-epidemiology (broad)0.0110.020
Bibliometrics0.0280.019
Science and technology studies0.0020.005
Scholarly communication0.0120.011
Open science0.0040.007
Research integrity0.0060.005
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.895
GPT teacher head0.688
Teacher spread0.207 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations24
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Evaluation in Clinical PracticeSame topicMeta-analysis and systematic reviewsFrench-language works237,207