Abstract 4143270: PREDICTIVE MODELS AID PHYSICIAN PROGNOSTICATION: A SECONDARY ANALYSIS EVALUATING INTEGRATED MODEL AND PHYSICIAN PROGNOSTIC ESTIMATES IN PATIENTS WITH HEART FAILURE WITH REDUCED EJECTION FRACTION
Bibliographic record
Abstract
Background: In recent studies from a multicenter Canadian cohort of outpatients with heart failure (HF), we found that model predictions were significantly more accurate than HF cardiologists. In this study, trying to mimic practice, we evaluated the additional predictive value and clinical impact of model predictions to refine physician estimated risk of 1-year mortality by combining model and physician estimates. Methods: We included consented consecutive HF outpatients (LVEF <40%) followed at 11 HF clinics in Canada. HF cardiologists estimated patient 1-year mortality using their clinical judgment. We calculated model predicted mortality using the Seattle HF Model (SHFM). We followed patients for at least a year to record mortality (or urgent heart transplant or ventricular assist device implant as mortality-equivalent events). Using random forest survival model and cross-validation, we compared the performance SHFM and the HF cardiologist alone, and the integrated HF cardiologist and the SHFM predictions by evaluating model discrimination (c-statistic), calibration (observed vs predicted event rate), risk reclassification and clinical net benefit analyses. Results: Among 1,643 HF patients, 1-year event rate was 9% (95%CI 8%-11%). The SHFM had the adequate discrimination (c-statistic 0.76) and excellent calibration while cardiologists showed adequate discrimination (c-statistic 0.75) and poor calibration with significant risk overestimation ( Figure 1 ). When the SHFM estimates were added physician predictions, discrimination significantly improved (0.82, 95%CI 0.78-0.86) with excellent calibration. By risk reclassification analysis, among patients with events, HF cardiologist better reclassified 44% than the SHFM or the integrated model. Among patients without event, however, HF cardiologists worse risk-classified 52% in comparison to SHFM and 71% to the integrated model. By net clinical benefit analysis ( Figure 2 ), when the decision to treat involves patients with 1-year mortality of >5%, SHFM predictions would lead to higher benefit than guiding care by physician judgement. Integrating model and HF cardiologist predictions led to minimally increased benefit in comparison to SHFM alone. Conclusions: Integrating prediction from the SHFM to physician judgment or using the SHFM alone showed superior accuracy than HF cardiologist predictions, proving that model-informed care may provide more accurate prognostic information to tailor clinical decision making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.031 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".