Systematic review of methods for individual patient data meta- analysis with binary outcomes
Bibliographic record
Abstract
BACKGROUND: Meta-analyses (MA) based on individual patient data (IPD) are regarded as the gold standard for meta-analyses and are becoming increasingly common, having several advantages over meta-analyses of summary statistics. These analyses are being undertaken in an increasing diversity of settings, often having a binary outcome. In a previous systematic review of articles published between 1999-2001, the statistical approach was seldom reported in sufficient detail, and the outcome was binary in 32% of the studies considered. Here, we explore statistical methods used for IPD-MA of binary outcomes only, a decade later. METHODS: We selected 56 articles, published in 2011 that presented results from an individual patient data meta-analysis. Of these, 26 considered a binary outcome. Here, we review 26 IPD-MA published during 2011 to consider: the goal of the study and reason for conducting an IPD-MA, whether they obtained all the data they sought, the approach used in their analysis, for instance, a two-stage or a one stage model, and the assumption of fixed or random effects. We also investigated how heterogeneity across studies was described and how studies investigated the effects of covariates. RESULTS: 19 of the 26 IPD-MA used a one-stage approach. 9 IPD-MA used a one-stage random treatment-effect logistic regression model, allowing the treatment effect to vary across studies. Twelve IPD-MA presented some form of statistic to measure heterogeneity across studies, though these were usually calculated using two-stage approach. Subgroup analyses were undertaken in all IPD-MA that aimed to estimate a treatment effect or safety of a treatment,. Sixteen meta-analyses obtained 90% or more of the patients sought. CONCLUSION: Evidence from this systematic review shows that the use of binary outcomes in assessing the effects of health care problems has increased, with random effects logistic regression the most common method of analysis. Methods are still often not reported in enough detail. Results also show that heterogeneity of treatment effects is discussed in most applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.093 | 0.336 |
| Meta-epidemiology (narrow) | 0.005 | 0.003 |
| Meta-epidemiology (broad) | 0.020 | 0.034 |
| Bibliometrics | 0.023 | 0.017 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.011 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".