Performance of current risk stratification models for predicting mortality in patients with heart failure: a systematic review and meta-analysis
Bibliographic record
Abstract
AIMS: There are several risk scores designed to predict mortality in patients with heart failure (HF). This study aimed to assess performance of risk scores validated for mortality prediction in patients with acute HF (AHF) and chronic HF. METHODS AND RESULTS: MEDLINE and Scopus were searched from January 2015 to January 2021 for studies which internally or externally validated risk models for predicting all-cause mortality in patients with AHF and chronic HF. Discrimination data were analysed using C-statistics, and pooled using generic inverse-variance random-effects model. Nineteen studies (n = 494 156 patients; AHF: 24 762; chronic HF mid-term mortality: 62 000; chronic HF long-term mortality: 452 097) and 11 risk scores were included. Overall, discrimination of risk scores was good across the three subgroups: AHF mortality [C-statistic: 0.76 (0.68-0.83)], chronic HF mid-term mortality [1 year; C-statistic: 0.74 (0.68-0.79)], and chronic HF long-term mortality [≥2 years; C-statistic: 0.71 (0.69-0.73)]. MEESSI-AHF [C-statistic: 0.81 (0.80-0.83)] and MARKER-HF [C-statistic: 0.85 (0.80-0.89)] had an excellent discrimination for AHF and chronic HF mid-term mortality, respectively, whereas MECKI had good discrimination [C-statistic: 0.78 (0.73-0.83)] for chronic HF long-term mortality relative to other models. Overall, risk scores predicting short-term mortality in patients with AHF did not have evidence of poor calibration (Hosmer-Lemeshow P > 0.05). However, risk models predicting mid-term and long-term mortality in patients with chronic HF varied in calibration performance. CONCLUSIONS: The majority of recently validated risk scores showed good discrimination for mortality in patients with HF. MEESSI-AHF demonstrated excellent discrimination in patients with AHF, and MARKER-HF and MECKI displayed an excellent discrimination in patients with chronic HF. However, modest reporting of calibration and lack of head-to-head comparisons in same populations warrant future studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.002 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".