Impact of Tumour Segmentation Accuracy on Efficacy of Quantitative MRI Biomarkers of Radiotherapy Outcome in Brain Metastasis
Bibliographic record
Abstract
Significantly affecting patients' clinical course and quality of life, a growing number of cancer cases are diagnosed with brain metastasis (BM) annually. Stereotactic radiotherapy is now a major treatment option for patients with BM. However, it may take months before the local response of BM to stereotactic radiation treatment is apparent on standard follow-up imaging. While machine learning in conjunction with radiomics has shown great promise in predicting the local response of BM before or early after radiotherapy, further development and widespread application of such techniques has been hindered by their dependency on manual tumour delineation. In this study, we explored the impact of using less-accurate automatically generated segmentation masks on the efficacy of radiomic features for radiotherapy outcome prediction in BM. The findings of this study demonstrate that while the effect of tumour delineation accuracy is substantial for segmentation models with lower dice scores (dice score ≤ 0.85), radiomic features and prediction models are rather resilient to imperfections in the produced tumour masks. Specifically, the selected radiomic features (six shared features out of seven) and performance of the prediction model (accuracy of 80% versus 80%, AUC of 0.81 versus 0.78) were fairly similar for the ground-truth and automatically generated segmentation masks, with dice scores close to 0.90. The positive outcome of this work paves the way for adopting high-throughput automatically generated tumour masks for discovering diagnostic and prognostic imaging biomarkers in BM without sacrificing accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".