Systematic review of methods for individual patient data meta- analysis with binary outcomes
Bibliographic record
Abstract
BACKGROUND: Meta-analyses (MA) based on individual patient data (IPD) are regarded as the gold standard for meta-analyses and are becoming increasingly common, having several advantages over meta-analyses of summary statistics. These analyses are being undertaken in an increasing diversity of settings, often having a binary outcome. In a previous systematic review of articles published between 1999-2001, the statistical approach was seldom reported in sufficient detail, and the outcome was binary in 32% of the studies considered. Here, we explore statistical methods used for IPD-MA of binary outcomes only, a decade later. METHODS: We selected 56 articles, published in 2011 that presented results from an individual patient data meta-analysis. Of these, 26 considered a binary outcome. Here, we review 26 IPD-MA published during 2011 to consider: the goal of the study and reason for conducting an IPD-MA, whether they obtained all the data they sought, the approach used in their analysis, for instance, a two-stage or a one stage model, and the assumption of fixed or random effects. We also investigated how heterogeneity across studies was described and how studies investigated the effects of covariates. RESULTS: 19 of the 26 IPD-MA used a one-stage approach. 9 IPD-MA used a one-stage random treatment-effect logistic regression model, allowing the treatment effect to vary across studies. Twelve IPD-MA presented some form of statistic to measure heterogeneity across studies, though these were usually calculated using two-stage approach. Subgroup analyses were undertaken in all IPD-MA that aimed to estimate a treatment effect or safety of a treatment,. Sixteen meta-analyses obtained 90% or more of the patients sought. CONCLUSION: Evidence from this systematic review shows that the use of binary outcomes in assessing the effects of health care problems has increased, with random effects logistic regression the most common method of analysis. Methods are still often not reported in enough detail. Results also show that heterogeneity of treatment effects is discussed in most applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.921 | 0.944 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.122 | 0.029 |
| Bibliometrics | 0.005 | 0.014 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.025 | 0.005 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.019 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".