MétaCan
Menu
Retour à la cohorte
Enregistrement W3155587303 · doi:10.1002/ejhf.2192

Cautious optimism for machine learning techniques for prediction of heart failure outcomes

2021· letter· en· W3155587303 sur OpenAlexaff
Nowell M. Fine, Jonathan G. Howlett

Notice bibliographique

RevueEuropean Journal of Heart Failure · 2021
Typeletter
Langueen
DomaineMedicine
ThématiqueHeart Failure Treatment and Management
Établissements canadiensLibin Cardiovascular Institute of AlbertaUniversity of Calgary
Organismes subventionnairesnon disponible
Mots-clésMachine learningArtificial intelligenceReceiver operating characteristicMedicinePredictive modellingHeart failurePredictive valueSet (abstract data type)Ejection fractionComputer scienceInternal medicine

Résumé

récupéré en direct d'OpenAlex

This article refers to 'A machine learning risk score predicts mortality across the spectrum of left ventricular ejection fraction' by B. Greenberg et al., published in this issue on pages 995–999. The desire to predict future events has been a basic human quality. It is therefore not surprising that numerous attempts to estimate mortality from heart failure (HF) have been made.1, 2 In most cases, a predictive model is constructed by performing one or more regression techniques upon a clinical dataset and validated in a different dataset. The predictive ability of any model is usually expressed as a c-statistic, referring to the area under a receiver operating characteristic curve (AUC), where a value of 1 denotes perfect predictive performance while a value of 0.5 denotes the performance of chance, such as with the flipping of a coin. To date, over 50 different models have been designed to predict HF mortality, incorporating as few as 5 and up to as many as 300 variables.3 These have reported moderate to good performance AUC values ranging from 0.6 to 0.8. Prediction of outcomes other than mortality has proven more challenging, with only fair to moderate predictive capability (0.55–0.7). These limitations have led to significant inertia in uptake of the use of predictive algorithms and to the study of alternative analytic techniques. Machine learning (ML), or the use of computer algorithms that allow for modification to identify complex patterns from large datasets, is a set of new techniques that are often categorized as artificial intelligence (AI).4 Application of these technologies has proliferated and evolved quickly across numerous fields of clinical care. HF is a complex and dynamic disorder with numerous multi-dimensional interactions and is associated with a high clinical and economic burden. In this context, risk prediction for HF outcomes would seem a particularly good fit for ML applications. However, early attempts to use ML-based prediction of HF outcomes initially demonstrated minimal incremental improvement over traditional methods,5, 6 which leads us to the present study. In a prior issue of this Journal, Adler et al.7 described the derivation and validation of a novel, ML-based algorithm for prediction of mortality in three distinct HF cohorts. To do this they employed two interesting techniques during model derivation. First, they excluded patients over the age of 80. Secondly, they identified a very high-risk group (those who died within 90 days) and a very low-risk group (those who did not die within 800 days). By comparing characteristics between these two vastly different groups, the authors built and trained a model using a boosted decision tree algorithm to relate subsets of the data to the two extreme outcomes. The resulting model employed eight common clinical/laboratory variables to predict all-cause mortality, generating a c-statistic ranging from 0.81–0.87 in three separate validation cohorts, which were superior to those obtained using other risk engines. Re-introduction of age to the model did not increase fidelity. In this issue of the Journal, Greenberg et al.8 extend their findings to demonstrate similar model performance irrespective of left ventricular ejection fraction (LVEF). To demonstrate this, they examined 4064 derivation cohort records with an echocardiographic measurement of LVEF taken within 30 days of study inception. Cases were categorized into one of three standard groups according to current HF guidelines9: HF with reduced ejection fraction (HFrEF; LVEF <40%, n = 782), HF with mid-range ejection fraction (HFmrEF; LVEF 40–49%, n = 404) and HF with preserved ejection fraction (HFpEF; LVEF ≥50%, n = 2878). For each group, Kaplan–Meier mortality curves were compared to predicted survival by log-rank test. The resulting c-statistics of 0.88, 0.83 and 0.85, respectively, were nearly identical to the overall model performance. Notably, the c-statistic for LVEF as a prediction of mortality in this dataset was only 0.52. While these findings are impressive, it is important to recognize the limitations inherent in the study design. There were relatively few patients in the HFmrEF and HFpEF groups, validation of the findings was not replicated in the validation cohorts, and special populations such as those without an available ejection fraction were not included. Hospitalization, a critical outcome closely related to cost and morbidity and that has to date remained resistant to accurate prediction, was not studied.5 Two key findings of this study should be emphasized: (i) the performance of this ML-based model irrespective of LVEF (a critical HF categorization) supports the internal consistency and hence further credence to its generalizability; (ii) the lack of independent predictive value of LVEF is consistent with previous work, although this consistency was not seen with age, in comparison to other models.10, 11 We are likely to observe surprising or even counterintuitive variable compositions in future ML-based predictive models. The future of ML-based prediction gives rise to several important considerations. AI-based modelling represents not one, but multiple different methods of modelling. The approach to algorithm development may vary with the clinical question or population outcome and may use rule-driven, decision tree or neural (deep learning) network-based algorithms, among others.12 In general, each technique is employed with care to balance the complexity of initial training steps with avoidance of over-fitting, which may lead to lack of generalizability. Another critical aspect is to use the 'best' data to develop models with the greatest utility. Even for models that demonstrate robust discrimination, inclusion of non-representative or incomplete data may lead to unanticipated performance, such as with underperformance of facial recognition in visible minority populations.13 Lack of relevant variables will necessarily affect performance. For instance, Sokoreli et al.14 demonstrated that by inclusion of patient-reported outcomes (items rarely collected in clinical systems), prediction of repeat hospitalization improved. With coalescence of regional electronic medical record, systems may introduce new heterogeneities that challenge widespread adoption of a single algorithm and favour region-specific algorithms. It cannot be emphasized strongly enough that given our incomplete understanding of the precise interaction between complex ML algorithms and the data they analyse, widespread validation should precede widespread usage of any predictive algorithm. With changes in our patients with HF and the treatments they receive, there will be a need to periodically review ML models to ensure ongoing validity. A final shared challenge is to determine how information gained from these AI-based algorithms is used, since the ultimate success of any tool relies on the decisions of those developing and using it. Conflict of interest: none declared.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,117
score de la tête « metaresearch » (Gemma)0,233
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Commentaire · Signal consensuel: Commentaire
Score de désaccord entre enseignants0,117
Score d'incertitude au seuil0,620

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,1170,233
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0030,003
Bibliométrie0,0040,003
Études des sciences et des technologies0,0030,023
Communication savante0,0160,025
Science ouverte0,0080,007
Intégrité de la recherche0,0120,070
Charge utile insuffisante (le modèle a refusé de juger)0,0060,008

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,024
Tête enseignante GPT0,266
Écart entre enseignants0,242 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreCommentaire

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations2
Publié2021
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueEuropean Journal of Heart FailureMême sujetHeart Failure Treatment and ManagementTravaux en français237 207