Evaluation of the replicability of systematic reviews with meta-analyses of the effects of health interventions
Bibliographic record
Abstract
ABSTRACT Background Systematic reviews are often characterised as being inherently replicable but several studies have challenged this claim. Objectives To investigate the variation in results following independent replication of literature searches and meta-analyses of systematic reviews. Methods We included ten systematic reviews of the effects of health interventions published in November 2020. Two information specialists repeated the original database search strategies. Two experienced review authors screened full-text articles, extracted data, and calculated the results for the first reported meta-analysis. All replicators were initially blinded to the results of the original review. A meta-analysis was considered not ‘fully replicable’ if the original and replicated summary estimate or confidence interval width differed by more than 10%, and meaningfully different if there was a difference in the direction or statistical significance. Results The difference between the number of records retrieved by the original reviewers and the information specialists exceeded 10% in 25/43 (58%) searches for the first replicator and 21/43 (49%) searches for the second. Eight meta-analyses (80%, 95% CI: 49-96%) were initially classified as not fully replicable. After screening and data discrepancies were addressed, the number of meta-analyses classified as not fully replicable decreased to five (50%, 95% CI: 24-76%). Differences were classified as meaningful in one blinded replication (10%, 95% CI: 1-40%) and none of the unblinded replications (0%, 95% CI: 0-28%). Conclusions The results of systematic review processes were not always consistent when their reported methods were repeated. However, these inconsistencies seldom affected summary estimates from meta-analyses in a meaningful way. HIGHLIGHTS What is already known on this topic Systematic reviews are often characterised as being inherently replicable, however, several studies have challenged this claim. Few studies have examined where and why inconsistencies arise, and what their impact is, when replicating multiple systematic review processes. What this study adds Replication of published systematic review processes (database searches, full-text screening, data extraction and meta-analysis) frequently produced results that were inconsistent with the original review. Following correction of replicator errors, the main drivers of variation in the results were incomplete reporting (e.g., unclear search methods, study eligibility criteria and methods for selecting study results) and reviewer data extraction errors. However, differences between the original reviewer’s and replicators’ summary estimates and confidence intervals were seldom meaningful.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.831 | 0.945 |
| Meta-epidemiology (narrow) | 0.007 | 0.006 |
| Meta-epidemiology (broad) | 0.020 | 0.047 |
| Bibliometrics | 0.018 | 0.019 |
| Science and technology studies | 0.003 | 0.016 |
| Scholarly communication | 0.010 | 0.016 |
| Open science | 0.012 | 0.013 |
| Research integrity | 0.010 | 0.009 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".