Evidence Synthesis for Complex Interventions Using Meta-Regression Models
Bibliographic record
Abstract
A goal of evidence synthesis for trials of complex interventions is to inform the design or implementation of novel versions of complex interventions by predicting expected outcomes with each intervention version. Conventional aggregate data meta-analyses of studies comparing complex interventions have limited ability to provide such information. We argue that evidence synthesis for trials of complex interventions should forgo aspirations of estimating causal effects and instead model the response surface of study results to 1) summarize the available evidence and 2) predict the average outcomes of future studies or in new settings. We illustrate this modeling approach using data from a systematic review of diabetes quality improvement (QI) interventions involving at least 1 of 12 QI strategy components. We specify a series of meta-regression models to assess the association of specific components with the posttreatment outcome mean and compare the results to conventional meta-analysis approaches. Compared with conventional approaches, modeling the response surface of study results can better reflect the associations between intervention components and study characteristics with the posttreatment outcome mean. Modeling study results using a response surface approach offers a useful and feasible goal for evidence synthesis of complex interventions that rely on aggregate data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.183 | 0.382 |
| Meta-epidemiology (narrow) | 0.006 | 0.003 |
| Meta-epidemiology (broad) | 0.016 | 0.044 |
| Bibliometrics | 0.017 | 0.013 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.006 |
| Open science | 0.006 | 0.005 |
| Research integrity | 0.004 | 0.006 |
| Insufficient payload (model declined to judge) | 0.011 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".