Integration of non-randomized studies with randomized controlled trials in meta-analyses of clinical studies: a meta-epidemiological study on effect estimation of interventions
Bibliographic record
Abstract
BACKGROUNDS: Syntheses of non-randomized studies of interventions (NRSIs) and randomized controlled trials (RCTs) are increasingly used in decision-making. This study aimed to summarize when NRSIs are included in evidence syntheses of RCTs, with a particular focus on the methodological issues associated with combining NRSIs and RCTs. METHODS: We searched PubMed to identify clinical systematic reviews published between 9 December 2017 and 9 December 2022, randomly sampling reviews in a 1:1 ratio of Core and non-Core clinical journals. We included systematic reviews with RCTs and NRSIs for the same clinical question. Clinical scenarios for considering the inclusion of NRSIs in eligible studies were classified. We extracted the methodological characteristics of the included studies, assessed the concordance of estimates between RCTs and NRSIs, calculated the ratio of the relative effect estimate from NRSIs to that from RCTs, and evaluated the impact on the estimates of pooled estimates when NRSIs are included. RESULTS: Two hundred twenty systematic reviews were included in the analysis. The clinical scenarios for including NRSIs were grouped into four main justifications: adverse outcomes (n = 140, 63.6%), long-term outcomes (n = 36, 16.4%), the applicability of RCT results to broader populations (n = 11, 5.0%), and other (n = 33, 15.0%). When conducting a meta-analysis, none of these reviews assessed the compatibility of the different types of evidence prior, 203 (92.3%) combined estimates from RCTs and NRSIs in the same meta-analysis. Of the 203 studies, 169 (76.8%) used crude estimates of NRSIs, and 28 (13.8%) combined RCTs and multiple types of NRSIs. Seventy-seven studies (35.5%) showed "qualitative disagree" between estimates from RCTs and NRSIs, and 101 studies (46.5%) found "important difference". The integration of NRSIs changed the qualitative direction of estimates from RCTs in 72 out of 200 studies (36.0%). CONCLUSIONS: Systematic reviews typically include NRSIs in the context of assessing adverse or long-term outcomes. The inclusion of NRSIs in a meta-analysis of RCTs has a substantial impact on effect estimates, but discrepancies between RCTs and NRSIs are often ignored. Our proposed recommendations will help researchers to consider carefully when and how to synthesis evidence from RCTs and NRSIs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchMeta-epidemiology (narrow)Meta-epidemiology (broad) Domain: Methods · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Meta-analysis | low |
| gpt | MetaresearchMeta-epidemiology (narrow)Meta-epidemiology (broad) Domain: Methods · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Meta-analysis | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.594 | 0.869 |
| Meta-epidemiology (narrow) | 0.008 | 0.007 |
| Meta-epidemiology (broad) | 0.033 | 0.075 |
| Bibliometrics | 0.040 | 0.040 |
| Science and technology studies | 0.002 | 0.005 |
| Scholarly communication | 0.018 | 0.019 |
| Open science | 0.008 | 0.012 |
| Research integrity | 0.009 | 0.009 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".