MétaCan
Menu
Back to cohort
Record W2103418324 · doi:10.1002/jrsm.53

Comparison of statistical inferences from the DerSimonian–Laird and alternative random‐effects model meta‐analyses – an empirical assessment of 920 Cochrane primary outcome meta‐analyses

2011· article· en· W2103418324 on OpenAlexaff
Kristian Thorlund, Jørn Wetterslev, Tahany Awad, Lehana Thabane, Christian Gluud

Bibliographic record

VenueResearch Synthesis Methods · 2011
Typearticle
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsSt. Joseph’s Healthcare HamiltonMcMaster University
Fundersnot available
KeywordsEstimatorMeta-analysisStatisticsRandom effects modelEconometricsMathematicsConfidence intervalMedicineInternal medicine

Abstract

fetched live from OpenAlex

In random-effects model meta-analysis, the conventional DerSimonian-Laird (DL) estimator typically underestimates the between-trial variance. Alternative variance estimators have been proposed to address this bias. This study aims to empirically compare statistical inferences from random-effects model meta-analyses on the basis of the DL estimator and four alternative estimators, as well as distributional assumptions (normal distribution and t-distribution) about the pooled intervention effect. We evaluated the discrepancies of p-values, 95% confidence intervals (CIs) in statistically significant meta-analyses, and the degree (percentage) of statistical heterogeneity (e.g. I(2)) across 920 Cochrane primary outcome meta-analyses. In total, 414 of the 920 meta-analyses were statistically significant with the DL meta-analysis, and 506 were not. Compared with the DL estimator, the four alternative estimators yielded p-values and CIs that could be interpreted as discordant in up to 11.6% or 6% of the included meta-analyses pending whether a normal distribution or a t-distribution of the intervention effect estimates were assumed. Large discrepancies were observed for the measures of degree of heterogeneity when comparing DL with each of the four alternative estimators. Estimating the degree (percentage) of heterogeneity on the basis of less biased between-trial variance estimators seems preferable to current practice. Disclosing inferential sensitivity of p-values and CIs may also be necessary when borderline significant results have substantial impact on the conclusion. Copyright © 2012 John Wiley & Sons, Ltd.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.411
metaresearch head score (Gemma)0.657
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Meta-analysis · Consensus signal: Meta-analysis
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.589
Threshold uncertainty score0.726

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.4110.657
Meta-epidemiology (narrow)0.0050.003
Meta-epidemiology (broad)0.0120.043
Bibliometrics0.0180.010
Science and technology studies0.0010.003
Scholarly communication0.0090.008
Open science0.0060.006
Research integrity0.0050.006
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.979
GPT teacher head0.786
Teacher spread0.193 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designMeta-analysis
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations80
Published2011
Admission routes1
Has abstractyes

Explore more

Same venueResearch Synthesis MethodsSame topicMeta-analysis and systematic reviewsFrench-language works237,207