Machine learning automated treatment planning for online magnetic resonance guided adaptive radiotherapy of prostate cancer
Bibliographic record
Abstract
Background and purpose: No best practices currently exist for achieving high quality radiation therapy (RT) treatment plan adaptation during magnetic resonance (MR) guided RT of prostate cancer. This study validates the use of machine learning (ML) automated RT treatment plan adaptation and benchmarks it against current clinical RT plan adaptation methods. Materials and methods: We trained an atlas-based ML automated treatment planning model using reference MR RT treatment plans (42.7 Gy in 7 fractions) from 46 patients with prostate cancer previously treated at our institution. For a held-out test set of 38 patients, retrospectively generated ML RT plans were compared to clinical human-generated adaptive RT plans for all 266 fractions. Differences in dose-volume metrics and clinical objective pass rates were evaluated using Wilcoxon tests (p < 0.05) and Exact McNemar tests (p < 0.05), respectively. Results: Compared to clinical RT plans, ML RT plans significantly increased sparing and objective pass rates of the rectum, bladder, and left femur. The mean ± standard deviation of rectum D20 and D50 in ML RT plans were 2.5 ± 2.2 Gy and 1.6 ± 1.3 Gy lower than clinical RT plans, respectively, with 14 % higher pass rates; bladder D40 was 4.6 ± 2.9 Gy lower with a 20 % higher pass rate; and the left femur D5 was 0.8 ± 1.8 Gy lower with a 7 % higher pass rate. Conclusions: ML automated RT treatment plan adaptation increases robustness to interfractional anatomical changes compared to current clinical adaptive RT practices by increasing compliance to treatment objectives.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".