65 Withdrawn
Bibliographic record
Abstract
Purpose: Stereotactic radiotherapy (SRT) for patients with brain metastasis requires precisely contoured gross tumour volumes (GTVs).We aimed to compare deep-learning-generated contours (DC) with expert contours (EC) on SRT treatment plans.Materials and Methods: Dataset 1 (DS1) consisted of 78 brain metastasis SRT plans with fused MPRAGE MR scans from a single centre.Dataset 2 (DS2) consisted of 170 publicly available MR images with brain metastasis contoured by a CNS radiation oncologist.Convolutional neural networks were used to train separate models on DS1, DS2, and a combined model including DS1 and DS2 (DS-all).Two model variations were developed: DC-Axial, which relied solely on axial training slices, and DC-Multi, which used axial contours combined with multiplanar slices in order to reduce false-positive predictions.A validation dataset consisted of 28 MPRAGE MR scans with 46 individual brain metastasis GTVs from SRT treatment plans.The true-positive DCs were compared with ECs using the Dice Similarity Coefficient (DSC), 95% distance transform (DT-95%), and mean distance transform (DT-mean).Dosimetric analysis was performed by fusing the DCs to the SRT planning CT scans with MIM Maestro (version 6.7).Results: The DC-Axial models identified 63%, 65%, 70% of metastasis for DS1, DS2, and DS-all respectively.The DC-Multi models identified 49%, 58%, 59% of metastasis for DS1, DS2, and DS-all respectively.The number of false positives for the DC-Axial models was 159, 138, 111 for DS1, DS2 and DS-all respectively.The number of false positives for the DC-Multi models was 12, nine, eight for DS1, DS2 and DS-all respectively.Comparing ECs to true-positive DS1, DS2 and DS-all models demonstrated a mean DSC of 0.68, 0.74, 0.77, a DT-mean of 0.77mm, 0.66mm, 0.57mm, and DT-95% of 1.71mm, 1.52 mm, 1.34mm respectively.The mean 80% isodose coverage was 100% for the ECs, 99.8% for DS1, 99.8% for DS2, and 99.9% for DS-all.The mean 90% isodose coverage was 100% for the ECs, 97% for DS1, 96% for DS2 and 97% for DS-all.There were no significant differences in the mean, max and minimum doses for DS1, DS2 and DS-all compared to the ECs. Conclusions:We observed accurate delineation of true-positive DCs on MPRAGE MR images which demonstrates the feasibility of using deep learning models to aid in tumour delineation for SRT treatment planning.Similar isodose coverage, and mean, max and minimum dose for the models further demonstrates the spatial agreement between the DCs and ECs.Models trained with the largest combined dataset (DS-all) had the best volumetric and dosimetric agreement with ECs.Multiplanar models demonstrated a significantly lower false-positive rate and slightly higher falsenegative rate for brain GTVs compared to models trained on axial slices alone.The true-positive detection rate can likely be improved in future studies that incorporate larger training datasets of MR images.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.006 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.004 | 0.005 |
| Insufficient payload (model declined to judge) | 0.419 | 0.344 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".