A systematic review and meta-analysis of sample size methodology for traumatic hemorrhage trials
Bibliographic record
Abstract
BACKGROUND: Trauma hemorrhage remains the most common cause of preventable mortality in trauma. To guide clinical practice, RCTs provide high-quality evidence to inform clinical decision making. The clinical relevance and inferences made by RCTs are dependent on assumptions made during sample size calculation. METHODS: To describe the quality of methodology for sample size determination, we conducted a systemic review RCTs evaluating interventions that aim to improve survival in adults with trauma-related hemorrhage. Estimated and actual outcome data are compared, including components of sample size determination. RESULTS: A total of 13 RCTs were included. We noted a high rate of negative trial results (11 of 13 studies). Most studies were multi-center and conducted in North America, evaluating patients with blunt and penetrating injuries. The criteria for hemorrhagic shock varied across studies. All studies did not accurately estimate the mortality rate during sample size calculation. All but one study overestimated the mortality reduction during sample size calculation; the median absolute mortality reduction was 3%, compared with a target of 10%. Only the CRASH-2 study used a minimal clinically important different for treatment effect target. No RCTs employed prognostic enrichment. Most studies were terminated (8 of 13), mainly for futility. CONCLUSION: Taken together, this review highlights that current clinical trial methodology is limited by imprecise control group risk estimates, overly optimistic treatment effect estimates, and lack of transparent justification for such targets. These limitations result in studies at high risk for futility and potentially premature abandonment of promising therapies. Given the high morbidity and mortality of trauma-related hemorrhage, we recommend that future conduct of trauma RCTs incorporate (1) prognostic enrichment to inform baseline risk, (2) justify target treatment differences based on clinical importance and realistic estimates of feasibility, and (3) be transparent and provide justification for the assumptions made. LEVEL OF EVIDENCE: Systematic Review/Meta-Analysis; Level III.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.042 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.017 | 0.006 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".