Knowledge-based automated planning with three-dimensional generative adversarial networks.
Bibliographic record
Abstract
PURPOSE: To develop a knowledge-based automated planning pipeline that generates treatment plans without feature engineering, using deep neural network architectures for predicting three-dimensional (3D) dose. METHODS: Our knowledge-based automated planning (KBAP) pipeline consisted of a knowledge-based planning (KBP) method that predicts dose for a contoured computed tomography (CT) image followed by two optimization models that learn objective function weights and generate fluence-based plans, respectively. We developed a novel generative adversarial network (GAN)-based KBP approach, a 3D GAN model, which predicts dose for the full 3D CT image at once and accounts for correlations between adjacent CT slices. Baseline comparisons were made against two state-of-the-art deep learning-based KBP methods from the literature. We also developed an additional benchmark, a two-dimensional (2D) GAN model which predicts dose to each axial slice independently. For all models, we investigated the impact of multiplicatively scaling the predictions before optimization, such that the predicted dose distributions achieved all target clinical criteria. Each KBP model was trained on 130 previously delivered oropharyngeal treatment plans. Performance was tested on 87 out-of-sample previously delivered treatment plans. All KBAP plans were evaluated using clinical planning criteria and compared to their corresponding clinical plans. KBP prediction quality was assessed using dose-volume histogram (DVH) differences from the corresponding clinical plans. RESULTS: The best performing KBAP plans were generated using predictions from the 3D GAN model that were multiplicatively scaled. These plans satisfied 77% of all clinical criteria, compared to the clinical plans, which satisfied 67% of all criteria. In general, multiplicatively scaling predictions prior to optimization increased the fraction of clinical criteria satisfaction by 11% relative to the plans generated with nonscaled predictions. Additionally, these KBAP plans satisfied the same criteria as the clinical plans 84% and 8% more frequently as compared to the two benchmark methods, respectively. CONCLUSIONS: We developed the first knowledge-based automated planning framework using a 3D generative adversarial network for prediction. Our results, based on 217 oropharyngeal cancer treatment plans, demonstrated superior performance in satisfying clinical criteria and generated more realistic plans as compared to the previous state-of-the-art approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".