MétaCan
Menu
← Back to cohort
Record W4412872054 · doi:10.1101/2025.07.17.25331623

Evaluation of the replicability of systematic reviews with meta-analyses of the effects of health interventions

2025· preprint· en· W4412872054 on OpenAlexaff
Daniel G. Hamilton, Joanne E. McKenzie, Phi‐Yen Nguyen, Melissa L. Rethlefsen, Steve McDonald, Sue Brennan, Fiona Fidler, Julian P. T. Higgins, Raju Kanukula, Sathya Karunananthan, Lara Maxwell, David Moher, Shinichi Nakagawa, David Nunan, Peter Tugwell, Vivian Welch, Matthew J. Page

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Languageen
FieldDecision Sciences
TopicMeta-analysis and systematic reviews
Canadian institutionsUniversity of AlbertaOttawa HospitalBruyèreUniversity of Ottawa
FundersNational Health and Medical Research CouncilMedical Research CouncilNational Institute for Health and Care Research
KeywordsPsychological interventionMeta-analysisSystematic reviewPsychologyMedicineMEDLINEPolitical sciencePsychiatryInternal medicine

Abstract

fetched live from OpenAlex

ABSTRACT Background Systematic reviews are often characterised as being inherently replicable but several studies have challenged this claim. Objectives To investigate the variation in results following independent replication of literature searches and meta-analyses of systematic reviews. Methods We included ten systematic reviews of the effects of health interventions published in November 2020. Two information specialists repeated the original database search strategies. Two experienced review authors screened full-text articles, extracted data, and calculated the results for the first reported meta-analysis. All replicators were initially blinded to the results of the original review. A meta-analysis was considered not ‘fully replicable’ if the original and replicated summary estimate or confidence interval width differed by more than 10%, and meaningfully different if there was a difference in the direction or statistical significance. Results The difference between the number of records retrieved by the original reviewers and the information specialists exceeded 10% in 25/43 (58%) searches for the first replicator and 21/43 (49%) searches for the second. Eight meta-analyses (80%, 95% CI: 49-96%) were initially classified as not fully replicable. After screening and data discrepancies were addressed, the number of meta-analyses classified as not fully replicable decreased to five (50%, 95% CI: 24-76%). Differences were classified as meaningful in one blinded replication (10%, 95% CI: 1-40%) and none of the unblinded replications (0%, 95% CI: 0-28%). Conclusions The results of systematic review processes were not always consistent when their reported methods were repeated. However, these inconsistencies seldom affected summary estimates from meta-analyses in a meaningful way. HIGHLIGHTS What is already known on this topic Systematic reviews are often characterised as being inherently replicable, however, several studies have challenged this claim. Few studies have examined where and why inconsistencies arise, and what their impact is, when replicating multiple systematic review processes. What this study adds Replication of published systematic review processes (database searches, full-text screening, data extraction and meta-analysis) frequently produced results that were inconsistent with the original review. Following correction of replicator errors, the main drivers of variation in the results were incomplete reporting (e.g., unclear search methods, study eligibility criteria and methods for selecting study results) and reviewer data extraction errors. However, differences between the original reviewer’s and replicators’ summary estimates and confidence intervals were seldom meaningful.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.831
metaresearch head score (Gemma)0.945
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.169
Threshold uncertainty score0.209

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.8310.945
Meta-epidemiology (narrow)0.0070.006
Meta-epidemiology (broad)0.0200.047
Bibliometrics0.0180.019
Science and technology studies0.0030.016
Scholarly communication0.0100.016
Open science0.0120.013
Research integrity0.0100.009
Insufficient payload (model declined to judge)0.0070.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.920
GPT teacher head0.639
Teacher spread0.281 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designObservational
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuemedRxiv→Same topicMeta-analysis and systematic reviews→French-language works237,207→