Machine learning for predictive analytics in medicine: real opportunity or overblown hype?
Notice bibliographique
Résumé
This editorial refers to ‘Predicting deterioration of ventricular function in patients with repaired tetralogy of Fallot using machine learning’ by M.D. Samad et al., pp. 730--738. Machine learning, a branch of computer sciences and an application of artificial intelligence (AI), is not new. As early as 1959, Samuel1 published in the IBM Journal the results of an experiment in which machine learning algorithms using rule-based systems to successfully learn the rules and strategy of checkers were developed that outperformed average players. However, despite early success and a relatively short existence, the field of AI has already experienced major cycles of disillusion, commonly known as AI winters. These winters have two things in common: astronomical expectations and subsequent failure to deliver anticipated results. Since the mid-1990s, AI has experienced renewed interest fuelled mainly by the increasing availability of computing power and the significant technological development and successful implementation of AI-based algorithms for decision support in many industries including banking, insurance, advertising, and transportation. The medical field shares this zeal, where interest in AI has skyrocketed following high-profile publications that showcase projects which have garnered significant media attention, and where the number of publications of AI-based projects in medical journals is steadily increasing. However, with this increased interest comes the risk of a new wave of overhyped expectations and a potential backlash should the technology fail to deliver; or deliver less than, or slower, than anticipated. And this risk is very real, in the 2017 Gartner Hype Cycle for Emerging Technologies deep learning was at the peak of inflated expectations,2 replacing the 2016 ‘winner’: machine learning.3 The risk of such backlash is particularly high in medicine given several factors. At the outset, there are astronomical, likely overinflated, expectations in regards to the current capabilities of AI algorithms. Additionally, there is a general lack of recognition of the technological and infrastructure gap between what AI algorithms need for successful implementation and the information technology infrastructure available in the majority of medical settings. Finally, and most importantly, there is an incomplete understanding on how to actually integrate the information generated by AI algorithms in the evidence-based paradigm of medicine. Recently, the high-profile failure of large AI implementation projects in various medical systems, demonstrate many of these hallmarks. And while some disillusion in regards to the potential of AI to transform medicine is probably inevitable; we must realign expectations, work diligently to bridge these gaps and most importantly rapidly integrate viable projects in health systems to show real-world feasibility and value. All of this will be necessary in order to prevent the same backlash that has previously vanquished so many promising technologies in the medical system.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,017 | 0,082 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,002 | 0,009 |
| Communication savante | 0,007 | 0,013 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,032 | 0,051 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,011 | 0,007 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».