When is a meta-analysis conclusive? A guide to Trial Sequential Analysis with an example of remote ischemic preconditioning for renoprotection in patients undergoing cardiac surgery
Bibliographic record
Abstract
Regardless of whether a randomized trial finds a statistically significant effect for an intervention or not, readers often wonder if the trial was large enough to be conclusive. To answer this question, we can estimate the required sample size for a trial by considering how commonly the outcome occurs, the smallest effect of clinical importance and the acceptable risk of falsely detecting or rejecting that effect. But when is a meta-analysis conclusive? We explain and illustrate the interpretation of Trial Sequential Analysis (TSA), a method increasingly used to answer this question. We conducted a conventional meta-analysis which suggested that, in adults undergoing cardiac surgery, remote ischemic preconditioning does not provide a statistically significant reduction in acute kidney injury (AKI) [12 trials, 4230 patients; relative risk 0.87 (95% confidence interval 0.74-1.02); P = 0.08; I2= 35%] or the risk of receiving acute dialysis [5 trials, 2111 patients; relative risk 1.15 (95% confidence interval 0.42-3.19); P = 0.78; I2 = 59%]. TSA demonstrates that as little as a 20% relative risk reduction in AKI is unlikely. Reliably finding effects on acute dialysis and smaller effects on AKI would require much more evidence. Notably, conventional meta-analyses conducted at one of the two earlier time points may have prematurely declared a statistically significant reduction in AKI, even though at no point in the TSA was there sufficient evidence to support such an effect. With this and other examples, we demonstrate that the TSA can prevent premature conclusions from meta-analyses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".