MétaCan
Menu
← Back to cohort
Record W4417013798 · doi:10.1182/blood-2025-4349

Evaluating large language models in real-world hematologic clinical decision-making: Performance, limitations, and clinical implications

2025· article· en· W4417013798 on OpenAlexaff
David M. Swoboda, Amy E. DeZern, James T. England, Sangeetha Venugopal, Thomas J. Kehoe, Brandon J. Aubrey, Marco Gabriele Raddi, Angela Consagra, Jiasheng Wang, Gustavo Rivero, Maximilian Stahl, Amer M. Zeidan, Torsten Haferlach, Andrew M. Brunner, Rena Buckstein, Valeria Santini, Matteo Giovanni Della Porta, Mikkael A. Sekeres, Aziz Nazha

Bibliographic record

VenueBlood · 2025
Typearticle
Languageen
FieldMedicine
TopicArtificial Intelligence in Healthcare and Education
Canadian institutionsHealth Sciences CentreSunnybrook Health Science Centre
Fundersnot available
KeywordsSubspecialtyHematologyHematologic NeoplasmsTest (biology)MEDLINEDiagnostic testSet (abstract data type)Scale (ratio)Specialty

Abstract

fetched live from OpenAlex

Abstract Background Recent advances in Artificial Intelligence (AI), particularly in Large Language Models (LLMs) like GPT-4o (ChatGPT) and others, have shown impressive performance in medical domains, including passing licensing exams and, in some cases, surpassing physicians in general diagnostic and reasoning tasks. However, their reliability and clinical utility in highly specialized, real-world medical settings—such as in hematology diagnostics and therapy—have not been rigorously evaluated. Malignant hematology poses unique challenges due to its complex pathophysiology, layered diagnostic frameworks, and the need for nuanced, high-stakes clinical decision-making mainly derived from highly specialized physicians, making it an ideal testbed for assessing the true capabilities and limitations of these models. Objectives To evaluate how well state-of-the-art LLMs handle real-world hematology cases, focusing on their ability to make accurate diagnoses, predict outcomes, follow treatment guidelines, and suggest relevant clinical trials. Method We developed a test set of 30 complex, real-world clinical cases of myelodysplastic syndromes (MDS). We chose MDS as a representative hematologic malignancy due to its diagnostic complexity and need for expert subspecialty care. Each case required integration of clinical, morphological, cytogenetic, and molecular data—mirroring real-life decision-making in hematology. A standardized prompt was used to query multiple LLMs—ChatGPT (GPT-4o and GPTo3), Claude, and DeepSeek. Models were tasked with providing a diagnosis per WHO 2022/ICC criteria, calculating IPSS-R/IPSS-M risk scores, and recommending appropriate treatment and clinical trials. Responses were independently reviewed by a blinded panel of eleven international MDS experts, who scored them on diagnostic accuracy, prognostic assessment, and treatment relevance on a Likert scale 1-5, with score ≥ 4 considered correct per expert opinion. Factual errors were also categorized as none, minor, or major. To evaluate the consistency of expert ratings, we used the intraclass correlation coefficient (ICC) to measure how well experts agreed on numerical scores and Cohen’s κ (kappa) to assess their agreement when identifying errors. Results The highest-performing model was GPT-o3, achieving 58% agreement with expert clinical assessments. This was followed by GPT-4o (42%), DeepSeek (31%), and Claude (26%). On a 1–5 scale, the average expert-assigned scores across domains were as follows: GPT-o3 (overall 3.48; Diagnosis 3.68, Prognosis 3.58, Treatment 3.56, Clinical Trials 3.09), GPT-4o (3.22; 3.14 / 3.20 / 3.39 / 3.16), DeepSeek (2.98; 2.92 / 3.01 / 3.15 / 2.83), and Claude (2.86; 2.72 / 2.90 / 3.09 / 2.73), respectively. Major factual errors (hallucinations) were frequent across all models, each exceeding a 25% rate: GPT-o3 and GPT-4o (both 26%), DeepSeek (33%), and Claude (36%). Minor factual error rates were similarly high: GPT-o3 and Claude (47%), DeepSeek (49%), and GPT-4o (52%). Experts showed strong agreement in their evaluations, with high consistency in scoring (ICC = 0.81) and in identifying AI errors or hallucinations (κ = 0.76), confirming the reliability of the review process. Conclusion Despite recent reports suggesting that models like ChatGPT have outperformed physicians in diagnostic accuracy and clinical decision-making, current state-of-the-art LLMs underperform in highly specialized and complex clinical scenarios in hematologic malignancies. Even advanced reasoning models such as GPT-o3 have fallen short of expert expectations. These findings underscore that general-purpose LLMs are not yet suitable for autonomous clinical use in hematology. Their deployment should be approached with caution, and further research is essential to rigorously evaluate their performance across all subdomains of hematology.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.017
metaresearch head score (Gemma)0.060
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.019
Threshold uncertainty score0.088

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0170.060
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0030.002
Open science0.0020.002
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.391
GPT teacher head0.583
Teacher spread0.193 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueBlood→Same topicArtificial Intelligence in Healthcare and Education→French-language works237,207→