RADT-10. THE LOST METASTASES: DEEP LEARNING’S POTENTIAL IN RADIOSURGERY QUALITY ASSURANCE
Bibliographic record
Abstract
Abstract Introduction Identifying, segmenting, measuring, and following multiple brain metastases treated with radiosurgery can be time consuming and error prone. Machine learning has shown promise for automated detection and segmentation. Recently, a U-Net inspired model combining volume aware loss functions and volume aware sampling methods was trained in an industrial-academic partnership. A total of 530 clinically annotated T1 gadolinium MRIs were used. Initial validation showed a high sensitivity (91%) with an average of 0.66 false positives per MRI. The goal of the present work was to characterize those “false positives” which may represent clinically undetected metastases. METHODS The images used for model development were clinically annotated for radiosurgery planning. Lesions had first been identified by a radiologist, second by clinicians during tumor board review, third by the treating radiation oncologist and the treating neurosurgeon (potentially after segmentation by a trainee) and finally by fellow radiation oncologists during quality assurance rounds. Despite these multiple checks, 10 patients (2%) had brain lesions considered potential clinical misses when all “false positives” were manually reviewed by a single investigator. Further detailed review including prior and subsequent imaging was used to arbitrate the nature of these lesions. RESULTS Among the 10 cases, four were confirmed as undetected metastases: two lesions required subsequent radiosurgery and 2 patients died prior to further imaging. The six other lesions were adjudicated as true “false positives” (typically vascular). CONCLUSION The multi-tier radiosurgery workflow at our institution left very few unidentified brain metastases (0.8%). Despite this low error rate, our AI algorithm still detected two lesions that required further treatment. Future investigations will focus on potential roles of AI in simplifying and accelerating our workflow. It also remains to be established if more undetected metastases would be seen in community settings where workflows include fewer sequential imaging reviews.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".