Do proton pump inhibitors increase mortality? A systematic review and in-depth analysis of the evidence
Bibliographic record
Abstract
Aims: Proton pump inhibitors (PPIs) were primarily approved for short term use (2 to 8 weeks). However, PPI use continues to expand. Widely believed to be safe, we reviewed emerging evidence on increased mortality with PPI long-term use. Methods: We searched MEDLINE, Embase and Cochrane Central for evidence from systematic reviews (SR) and primary studies reporting all-cause mortality in adults treated with a PPI for any indication (duration > 12 weeks) compared to patients without PPI treatment (no use, placebo or H2RA use). Data was synthesized, analysed, critically examined and interpreted herein. Results: From 1304 articles, one systematic review (SR) was identified that reported on all-cause mortality. The SR pooled 3 observational studies with data to 1 year: odds ratio, 95% confidence interval (CI) 1.53-1.84. A randomized controlled trial (RCT), the COMPASS (Cardiovascular Outcomes for People Using Anticoagulant Strategies) RCT with data to 3 years: hazard ratio (HR) 1.03, 95% CI 0.92-1.15. The US Veterans Affairs cohort study using a large national dataset with data to 10 years; HR 1.17, 95% CI (1.10-1.24), (NNH) 22. The most common causes of death were from cardiovascular and chronic kidney diseases, with an excess death of 15 and 4 per 1000 patients, respectively over 10-year period. Conclusions: Harms arising from real world medication use are best evaluated using a pharmacovigilance ‘convergence of proof’ approach using data from a variety of sources and varied study designs. Careful appraisal of the totality of available evidence leads to the conclusion that long-term PPI utilization increases mortality
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.000 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".