MétaCan
Menu
← Retour à la cohorte
Enregistrement W4417013798 · doi:10.1182/blood-2025-4349

Evaluating large language models in real-world hematologic clinical decision-making: Performance, limitations, and clinical implications

2025· article· en· W4417013798 sur OpenAlexaff
David M. Swoboda, Amy E. DeZern, James T. England, Sangeetha Venugopal, Thomas J. Kehoe, Brandon J. Aubrey, Marco Gabriele Raddi, Angela Consagra, Jiasheng Wang, Gustavo Rivero, Maximilian Stahl, Amer M. Zeidan, Torsten Haferlach, Andrew M. Brunner, Rena Buckstein, Valeria Santini, Matteo Giovanni Della Porta, Mikkael A. Sekeres, Aziz Nazha

Notice bibliographique

RevueBlood · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueArtificial Intelligence in Healthcare and Education
Établissements canadiensHealth Sciences CentreSunnybrook Health Science Centre
Organismes subventionnairesnon disponible
Mots-clésSubspecialtyHematologyHematologic NeoplasmsTest (biology)MEDLINEDiagnostic testSet (abstract data type)Scale (ratio)Specialty

Résumé

récupéré en direct d'OpenAlex

Abstract Background Recent advances in Artificial Intelligence (AI), particularly in Large Language Models (LLMs) like GPT-4o (ChatGPT) and others, have shown impressive performance in medical domains, including passing licensing exams and, in some cases, surpassing physicians in general diagnostic and reasoning tasks. However, their reliability and clinical utility in highly specialized, real-world medical settings—such as in hematology diagnostics and therapy—have not been rigorously evaluated. Malignant hematology poses unique challenges due to its complex pathophysiology, layered diagnostic frameworks, and the need for nuanced, high-stakes clinical decision-making mainly derived from highly specialized physicians, making it an ideal testbed for assessing the true capabilities and limitations of these models. Objectives To evaluate how well state-of-the-art LLMs handle real-world hematology cases, focusing on their ability to make accurate diagnoses, predict outcomes, follow treatment guidelines, and suggest relevant clinical trials. Method We developed a test set of 30 complex, real-world clinical cases of myelodysplastic syndromes (MDS). We chose MDS as a representative hematologic malignancy due to its diagnostic complexity and need for expert subspecialty care. Each case required integration of clinical, morphological, cytogenetic, and molecular data—mirroring real-life decision-making in hematology. A standardized prompt was used to query multiple LLMs—ChatGPT (GPT-4o and GPTo3), Claude, and DeepSeek. Models were tasked with providing a diagnosis per WHO 2022/ICC criteria, calculating IPSS-R/IPSS-M risk scores, and recommending appropriate treatment and clinical trials. Responses were independently reviewed by a blinded panel of eleven international MDS experts, who scored them on diagnostic accuracy, prognostic assessment, and treatment relevance on a Likert scale 1-5, with score ≥ 4 considered correct per expert opinion. Factual errors were also categorized as none, minor, or major. To evaluate the consistency of expert ratings, we used the intraclass correlation coefficient (ICC) to measure how well experts agreed on numerical scores and Cohen’s κ (kappa) to assess their agreement when identifying errors. Results The highest-performing model was GPT-o3, achieving 58% agreement with expert clinical assessments. This was followed by GPT-4o (42%), DeepSeek (31%), and Claude (26%). On a 1–5 scale, the average expert-assigned scores across domains were as follows: GPT-o3 (overall 3.48; Diagnosis 3.68, Prognosis 3.58, Treatment 3.56, Clinical Trials 3.09), GPT-4o (3.22; 3.14 / 3.20 / 3.39 / 3.16), DeepSeek (2.98; 2.92 / 3.01 / 3.15 / 2.83), and Claude (2.86; 2.72 / 2.90 / 3.09 / 2.73), respectively. Major factual errors (hallucinations) were frequent across all models, each exceeding a 25% rate: GPT-o3 and GPT-4o (both 26%), DeepSeek (33%), and Claude (36%). Minor factual error rates were similarly high: GPT-o3 and Claude (47%), DeepSeek (49%), and GPT-4o (52%). Experts showed strong agreement in their evaluations, with high consistency in scoring (ICC = 0.81) and in identifying AI errors or hallucinations (κ = 0.76), confirming the reliability of the review process. Conclusion Despite recent reports suggesting that models like ChatGPT have outperformed physicians in diagnostic accuracy and clinical decision-making, current state-of-the-art LLMs underperform in highly specialized and complex clinical scenarios in hematologic malignancies. Even advanced reasoning models such as GPT-o3 have fallen short of expert expectations. These findings underscore that general-purpose LLMs are not yet suitable for autonomous clinical use in hematology. Their deployment should be approached with caution, and further research is essential to rigorously evaluate their performance across all subdomains of hematology.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,017
score de la tête « metaresearch » (Gemma)0,060
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,019
Score d'incertitude au seuil0,088

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0170,060
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,001
Études des sciences et des technologies0,0010,001
Communication savante0,0030,002
Science ouverte0,0020,002
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,0030,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,391
Tête enseignante GPT0,583
Écart entre enseignants0,193 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueBlood→Même sujetArtificial Intelligence in Healthcare and Education→Travaux en français237 207→