MétaCan
Menu
Back to cohort
Record W3155587303 · doi:10.1002/ejhf.2192

Cautious optimism for machine learning techniques for prediction of heart failure outcomes

2021· letter· en· W3155587303 on OpenAlexaff
Nowell M. Fine, Jonathan G. Howlett

Bibliographic record

VenueEuropean Journal of Heart Failure · 2021
Typeletter
Languageen
FieldMedicine
TopicHeart Failure Treatment and Management
Canadian institutionsLibin Cardiovascular Institute of AlbertaUniversity of Calgary
Fundersnot available
KeywordsMachine learningArtificial intelligenceReceiver operating characteristicMedicinePredictive modellingHeart failurePredictive valueSet (abstract data type)Ejection fractionComputer scienceInternal medicine

Abstract

fetched live from OpenAlex

This article refers to 'A machine learning risk score predicts mortality across the spectrum of left ventricular ejection fraction' by B. Greenberg et al., published in this issue on pages 995–999. The desire to predict future events has been a basic human quality. It is therefore not surprising that numerous attempts to estimate mortality from heart failure (HF) have been made.1, 2 In most cases, a predictive model is constructed by performing one or more regression techniques upon a clinical dataset and validated in a different dataset. The predictive ability of any model is usually expressed as a c-statistic, referring to the area under a receiver operating characteristic curve (AUC), where a value of 1 denotes perfect predictive performance while a value of 0.5 denotes the performance of chance, such as with the flipping of a coin. To date, over 50 different models have been designed to predict HF mortality, incorporating as few as 5 and up to as many as 300 variables.3 These have reported moderate to good performance AUC values ranging from 0.6 to 0.8. Prediction of outcomes other than mortality has proven more challenging, with only fair to moderate predictive capability (0.55–0.7). These limitations have led to significant inertia in uptake of the use of predictive algorithms and to the study of alternative analytic techniques. Machine learning (ML), or the use of computer algorithms that allow for modification to identify complex patterns from large datasets, is a set of new techniques that are often categorized as artificial intelligence (AI).4 Application of these technologies has proliferated and evolved quickly across numerous fields of clinical care. HF is a complex and dynamic disorder with numerous multi-dimensional interactions and is associated with a high clinical and economic burden. In this context, risk prediction for HF outcomes would seem a particularly good fit for ML applications. However, early attempts to use ML-based prediction of HF outcomes initially demonstrated minimal incremental improvement over traditional methods,5, 6 which leads us to the present study. In a prior issue of this Journal, Adler et al.7 described the derivation and validation of a novel, ML-based algorithm for prediction of mortality in three distinct HF cohorts. To do this they employed two interesting techniques during model derivation. First, they excluded patients over the age of 80. Secondly, they identified a very high-risk group (those who died within 90 days) and a very low-risk group (those who did not die within 800 days). By comparing characteristics between these two vastly different groups, the authors built and trained a model using a boosted decision tree algorithm to relate subsets of the data to the two extreme outcomes. The resulting model employed eight common clinical/laboratory variables to predict all-cause mortality, generating a c-statistic ranging from 0.81–0.87 in three separate validation cohorts, which were superior to those obtained using other risk engines. Re-introduction of age to the model did not increase fidelity. In this issue of the Journal, Greenberg et al.8 extend their findings to demonstrate similar model performance irrespective of left ventricular ejection fraction (LVEF). To demonstrate this, they examined 4064 derivation cohort records with an echocardiographic measurement of LVEF taken within 30 days of study inception. Cases were categorized into one of three standard groups according to current HF guidelines9: HF with reduced ejection fraction (HFrEF; LVEF <40%, n = 782), HF with mid-range ejection fraction (HFmrEF; LVEF 40–49%, n = 404) and HF with preserved ejection fraction (HFpEF; LVEF ≥50%, n = 2878). For each group, Kaplan–Meier mortality curves were compared to predicted survival by log-rank test. The resulting c-statistics of 0.88, 0.83 and 0.85, respectively, were nearly identical to the overall model performance. Notably, the c-statistic for LVEF as a prediction of mortality in this dataset was only 0.52. While these findings are impressive, it is important to recognize the limitations inherent in the study design. There were relatively few patients in the HFmrEF and HFpEF groups, validation of the findings was not replicated in the validation cohorts, and special populations such as those without an available ejection fraction were not included. Hospitalization, a critical outcome closely related to cost and morbidity and that has to date remained resistant to accurate prediction, was not studied.5 Two key findings of this study should be emphasized: (i) the performance of this ML-based model irrespective of LVEF (a critical HF categorization) supports the internal consistency and hence further credence to its generalizability; (ii) the lack of independent predictive value of LVEF is consistent with previous work, although this consistency was not seen with age, in comparison to other models.10, 11 We are likely to observe surprising or even counterintuitive variable compositions in future ML-based predictive models. The future of ML-based prediction gives rise to several important considerations. AI-based modelling represents not one, but multiple different methods of modelling. The approach to algorithm development may vary with the clinical question or population outcome and may use rule-driven, decision tree or neural (deep learning) network-based algorithms, among others.12 In general, each technique is employed with care to balance the complexity of initial training steps with avoidance of over-fitting, which may lead to lack of generalizability. Another critical aspect is to use the 'best' data to develop models with the greatest utility. Even for models that demonstrate robust discrimination, inclusion of non-representative or incomplete data may lead to unanticipated performance, such as with underperformance of facial recognition in visible minority populations.13 Lack of relevant variables will necessarily affect performance. For instance, Sokoreli et al.14 demonstrated that by inclusion of patient-reported outcomes (items rarely collected in clinical systems), prediction of repeat hospitalization improved. With coalescence of regional electronic medical record, systems may introduce new heterogeneities that challenge widespread adoption of a single algorithm and favour region-specific algorithms. It cannot be emphasized strongly enough that given our incomplete understanding of the precise interaction between complex ML algorithms and the data they analyse, widespread validation should precede widespread usage of any predictive algorithm. With changes in our patients with HF and the treatments they receive, there will be a need to periodically review ML models to ensure ongoing validity. A final shared challenge is to determine how information gained from these AI-based algorithms is used, since the ultimate success of any tool relies on the decisions of those developing and using it. Conflict of interest: none declared.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.117
metaresearch head score (Gemma)0.233
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.117
Threshold uncertainty score0.620

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1170.233
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0040.003
Science and technology studies0.0030.023
Scholarly communication0.0160.025
Open science0.0080.007
Research integrity0.0120.070
Insufficient payload (model declined to judge)0.0060.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.024
GPT teacher head0.266
Teacher spread0.242 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueEuropean Journal of Heart FailureSame topicHeart Failure Treatment and ManagementFrench-language works237,207