Can machine learning models improve early detection of brain metastases using diffusion weighted imaging-based radiomics?
Bibliographic record
Abstract
Background: Metastatic complications are a major cause of cancer-related morbidity, with up to 40% of cancer patients experiencing at least one brain metastasis. Earlier detection may significantly improve patient outcomes and overall survival. We investigated machine learning (ML) models for early detection of brain metastases based on diffusion weighted imaging (DWI) radiomics. Methods: Longitudinal diffusion imaging from 116 patients previously treated with stereotactic radiosurgery (SRS) for brain metastases were retrospectively analyzed. Clinical contours from 600 metastases were extracted from radiosurgery planning computed tomography, and rigidly registered to corresponding contrast enhanced-T1 and apparent diffusion coefficient (ADC) maps. Contralateral contours located in healthy brain tissue were used as control. The dataset consisted of (I) radiomic features using ADC maps, (II) radiomic feature change calculated using timepoints before the metastasis manifested on contrast enhanced-T1, (III) primary cancer, and (IV) anatomical location. The dataset was divided into training and internal validation sets using an 80/20 split with stratification. Four classification algorithms [Linear Support Vector Machine (SVM), Random Forest (RF), AdaBoost, and XGBoost] underwent supervised classification training, with contours labeled either 'control' or 'metastasis'. Hyperparameters were optimized towards balanced accuracy. Various model metrics (receiver operating characteristic curve area scores, accuracy, recall, and precision) were calculated to gauge performance. Results: The radiomic and clinical data set, feature engineering, and ML models developed were able to identify metastases with an accuracy of up to 87.7% on the training set, and 85.8% on an unseen test set. XGBoost and RF showed superior accuracy (XGBoost: 0.877±0.021 and 0.833±0.47, RF: 0.823±0.024 and 0.858±0.045) for training and validation sets, respectively. XGBoost and RF also showed strong area under the receiver operating characteristic curve (AUC) performance on the validation set (0.910±0.037 and 0.922±0.034, respectively). AdaBoost performed slightly lower in all metrics. SVM model generalized poorly with the internal validation set. Important features involved changes in radiomics months before manifesting on contrast enhanced-T1. Conclusions: The proposed models using diffusion-based radiomics showed encouraging results in differentiating healthy brain tissue from metastases using clinical imaging data. These findings suggest that longitudinal diffusion imaging and ML may help improve patient care through earlier diagnosis and increased patient monitoring/follow-up. Future work aims to improve model classification metrics, robustness, user-interface, and clinical applicability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".