Can machine learning models improve early detection of brain metastases using diffusion weighted imaging-based radiomics?
Bibliographic record
Abstract
Background: Metastatic complications are a major cause of cancer-related morbidity, with up to 40% of cancer patients experiencing at least one brain metastasis. Earlier detection may significantly improve patient outcomes and overall survival. We investigated machine learning (ML) models for early detection of brain metastases based on diffusion weighted imaging (DWI) radiomics. Methods: Longitudinal diffusion imaging from 116 patients previously treated with stereotactic radiosurgery (SRS) for brain metastases were retrospectively analyzed. Clinical contours from 600 metastases were extracted from radiosurgery planning computed tomography, and rigidly registered to corresponding contrast enhanced-T1 and apparent diffusion coefficient (ADC) maps. Contralateral contours located in healthy brain tissue were used as control. The dataset consisted of (I) radiomic features using ADC maps, (II) radiomic feature change calculated using timepoints before the metastasis manifested on contrast enhanced-T1, (III) primary cancer, and (IV) anatomical location. The dataset was divided into training and internal validation sets using an 80/20 split with stratification. Four classification algorithms [Linear Support Vector Machine (SVM), Random Forest (RF), AdaBoost, and XGBoost] underwent supervised classification training, with contours labeled either 'control' or 'metastasis'. Hyperparameters were optimized towards balanced accuracy. Various model metrics (receiver operating characteristic curve area scores, accuracy, recall, and precision) were calculated to gauge performance. Results: The radiomic and clinical data set, feature engineering, and ML models developed were able to identify metastases with an accuracy of up to 87.7% on the training set, and 85.8% on an unseen test set. XGBoost and RF showed superior accuracy (XGBoost: 0.877±0.021 and 0.833±0.47, RF: 0.823±0.024 and 0.858±0.045) for training and validation sets, respectively. XGBoost and RF also showed strong area under the receiver operating characteristic curve (AUC) performance on the validation set (0.910±0.037 and 0.922±0.034, respectively). AdaBoost performed slightly lower in all metrics. SVM model generalized poorly with the internal validation set. Important features involved changes in radiomics months before manifesting on contrast enhanced-T1. Conclusions: The proposed models using diffusion-based radiomics showed encouraging results in differentiating healthy brain tissue from metastases using clinical imaging data. These findings suggest that longitudinal diffusion imaging and ML may help improve patient care through earlier diagnosis and increased patient monitoring/follow-up. Future work aims to improve model classification metrics, robustness, user-interface, and clinical applicability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".