MétaCan
Menu
Back to cohort
Record W4384155640 · doi:10.1101/2023.07.13.23292601

Design differences explain variation in results between randomized trials and their non-randomized emulations

2023· preprint· en· W4384155640 on OpenAlexfundno aff
Rachel Heyard, Leonhard Held, Sebastian Schneeweiß, Shirley Wang

Bibliographic record

VenuemedRxiv · 2023
Typepreprint
Languageen
FieldMathematics
TopicAdvanced Causal Inference Techniques
Canadian institutionsnot available
FundersU.S. Food and Drug AdministrationHamilton Health Sciences FoundationNational Institutes of HealthBrigham and Women's Hospital
KeywordsRandomized controlled trialMedicineConfoundingDiscontinuationRandomized experimentResearch designSample size determinationCompletely randomized designRandomizationObservational studyStatisticsSurgeryInternal medicine

Abstract

fetched live from OpenAlex

ABSTRACT Objectives While randomized controlled trials (RCTs) are considered a standard for evidence on the efficacy of medical treatments, non-randomized real-world evidence (RWE) studies using data from health insurance claims or electronic health records can provide important complementary evidence. The use of RWE to inform decision-making has been questioned because of concerns regarding confounding in non-randomized studies and the use of secondary data. RCT-DUPLICATE was a demonstration project that emulated the design of 32 RCTs with non-randomized RWE studies. We sought to explore how emulation differences relate to variation in results between the RCT-RWE study pairs. Methods We include all RCT-RWE study pairs from RCT-DUPLICATE where the measure of effect was a hazard ratio and use exploratory meta-regression methods to explain differences and variation in the effect sizes between the results from the RCT and the RWE study. The considered explanatory variables are related to design and population differences. Results Most of the observed variation in effect estimates between RCT-RWE study pairs in this sample could be explained by three emulation differences in the meta-regression model: (i) in-hospital start of treatment (not observed in claims data), (ii) discontinuation of certain baseline therapies at randomization (not part of clinical practice), (iii) delayed onset of drug effects (missed by short medication persistence in clinical practice). Conclusions This analysis suggests that a substantial proportion of the observed variation between results from RCTs and RWE studies can be attributed to design emulation differences. What is already known on this topic Real-world evidence (RWE) studies can complement randomized controlled trials (RCT) by providing insights on the effectiveness of a medical treatment in clinical practice. Concerns about confounding have limited the use of RWE studies in clinical practice and policy decisions. What this study adds A large share of the observed variation in results between RCT-RWE study pairs could be explained by design emulation differences.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.591
metaresearch head score (Gemma)0.777
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Methods · Consensus signal: Methods
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.409
Threshold uncertainty score0.504

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.5910.777
Meta-epidemiology (narrow)0.0040.003
Meta-epidemiology (broad)0.0070.024
Bibliometrics0.0060.006
Science and technology studies0.0010.008
Scholarly communication0.0070.006
Open science0.0060.006
Research integrity0.0050.005
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.426
GPT teacher head0.443
Teacher spread0.017 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designSimulation or modeling
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2023
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicAdvanced Causal Inference TechniquesFrench-language works237,207