Individual participant data meta-analyses compared with meta-analyses based on aggregate data
Bibliographic record
Abstract
Meta-analysis based on individual participant data (IPD) is widely accepted as the most reliable approach and has been described as the ‘gold standard‘ for systematic reviews. An IPD approach often allows more powerful, consistent and thorough analyses but does require additional resources compared with meta-analyses based on aggregate data (AD). Several empirical comparisons of IPD with AD meta-analyses have been published, some of which show that IPD meta-analyses can differ in important ways from meta-analyses based on AD. For example, the importance of including as much follow-up as possible on all randomised participants and data from all relevant trials was shown in separate empirical comparisons [ 1 – 3 ] whilst Duchateau [ 4 ] found substantial differences between IPD and literature-based AD meta-analyses, mainly due to different approaches to analysis. An unpublished review [ 5 ] summarised results from across 25 studies, showing that for two thirds of the comparisons AD estimated effect sizes with less precision and tended to overestimate the IPD effect but differences were small in most comparisons. We have undertaken a Cochrane systematic review of empirical studies that compared IPD and AD meta-analysis to explore key reasons for the differences. The Cochrane Methodology Register, CENTRAL, MEDLINE and EMBASE were searched using a predefined set of search terms. Studies that report an empirical comparison of IPD meta-analysis against AD meta-analysis of randomised trials were assessed for inclusion, by two reviewers independently. Data were extracted by two reviewers independently and stored in a central database. Over forty empirical studies have satisfied the inclusion criteria for this review. Estimates of effect size and precision obtained from IPD and AD will be compared and differences will be discussed. Results will help inform the ongoing debate about whether, and when IPD may be most valuable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".