MétaCan
Menu
← Back to cohort

Abstract A047: Building clinically relevant and interpretable multi-modal machine learning algorithms to predict glioblastoma disease progression

2025· article· en· W4412163894 on OpenAlexaboutno aff
Shreya Chappidi, Erdal Taşçı, Longze Zhang, Kevin Camphausen, Andra Krauze

Bibliographic record

VenueClinical Cancer Research · 2025
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsnot available
Fundersnot available
KeywordsGlioblastomaModalAlgorithmDiseaseMedicineArtificial intelligenceComputer scienceMachine learningPathologyCancer research

Abstract

fetched live from OpenAlex

Abstract Introduction: Machine learning (ML) algorithms have demonstrated high predictive performance in clinical diagnostics and decision making. However, these applications often suffer from low clinical adoption due to poor interpretability and unimodal workflows (e.g., trained only on imaging) that do not mirror real-world clinical reasoning. Predicting disease progression is essential for early management and improved outcomes in glioblastoma (GBM), a highly recurrent primary brain tumor with poor prognosis. We develop a novel multi-modal ML algorithm for GBM progression prediction using computer vision (CV) and natural language processing (NLP) algorithms to extract interpretable, clinically relevant features for further downstream outcome prediction. Methods: In a cohort of 92 pathology-proven GBM patients, we analyzed patient data including free text notes, MRI imaging, radiotherapy (RT) dosimetry, medications, and clinical variables. The target prediction was 24-month progression-free survival (PFS), with ground truth manually obtained via expert application of Response Assessment in Neuro-Oncology (RANO) clinical criteria (median 15.3 mo). Our custom local spaCy NLP model extracted RANO terms relating to disease progression and stability, while deepMedic tumor segmentation CV algorithms quantified tumor volume changes in brain MRI scans. A random forest model was trained on term frequencies, tumor volume changes, and scaled clinical data using an 80/20 train-test split with stratification to handle class imbalance. Results: Our multi-modal algorithm with clinically interpretable features yielded a 95% training accuracy and test performance metrics of 90% accuracy, 94% F1 score, and 91% AUROC. Feature importance obtained via mean decrease in impurity indicated that the top five features influencing model performance included max RT dose delivered to the optic chiasm and brain stem, frequency of stability-related radiology report terms, age at diagnosis, and 6-month change in contrast-enhancing tumor volume. These results highlight the potential utility of mirroring clinical progression criteria via intermediate NLP- and CV-derived features and using RT dosimetry parameters, which are underemployed in cancer ML applications. Conclusion: This study reports a novel multi-modal ML approach to GBM progression prediction using interpretable intermediate clinical features extracted from raw imaging, RT dosimetry, and free text data. Further user studies with clinicians and larger patient cohorts can validate and fine-tune this interpretable multi-modal approach to approximating real-world clinical reasoning and contrast results with current black box, unimodal algorithms. Citation Format: Shreya Chappidi, Erdal Tasci, Longze Zhang, Kevin Camphausen, Andra V. Krauze. Building clinically relevant and interpretable multi-modal machine learning algorithms to predict glioblastoma disease progression [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A047.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.006
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.007
Threshold uncertainty score0.014

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.006
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.000
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.062
GPT teacher head0.517
Teacher spread0.455 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueClinical Cancer Research→Same topicRadiomics and Machine Learning in Medical Imaging→French-language works237,207→