Multi-modality Artificial Intelligence for Involved-Site Radiation Therapy: Clinical Target Volume Delineation in High-Risk Pediatric Hodgkin Lymphoma
Bibliographic record
Abstract
Introduction Clinical target volume (CTV) delineation for involved-site radiation therapy (ISRT) in Hodgkin lymphoma (HL) is time-consuming due to the need to analyze multi-time-point PET/CT scans co-registered to the planning CT. Deep learning (DL) has the potential to streamline this task, but its feasibility remains unexplored. Our goal was to develop automated CTV segmentation algorithms that integrated multi-modality imaging to facilitate ISRT planning. Methods This study included planning CT, baseline PET/CT (PET1), and interim PET/CT (PET2) scans from 288 pediatric patients with high-risk HL in the Children’s Oncology Group AHOD 1331 trial. Data from 58 patients across 24 institutions were held out for external testing, while the remaining 230 cases from 95 institutions were used for model development. We investigated three DL architectures (SegResNet, ResUNet, and SwinUNETR) and evaluated the impact of incorporating PET1 and PET2 images alongside the planning CT. Performance was assessed using the 95th percentile Hausdorff distance (HD95) and Dice similarity coefficient (DSC). Inter-observer variability (IOV) was estimated by comparing original institutional CTVs with those newly delineated by four board-certified radiation oncologists on a subset of 10 cases. The quality of CTVs generated by the top-performing model and those from original institutions was independently assessed on 40 other cases by four radiation oncologists, who were blinded to the source of the CTVs. Results On the external cohort, a SwinUNETR model incorporating planning CT, PET1, and PET2 images achieved the highest performance, with an HD95 of 34.43 mm, and DSC of 0.72. In comparison, the best planning CT-only model attained an HD95 of 58.94 mm and DSC of 0.68. All models incorporating PET/CT images were significantly better (P<0.01) than CT-only models. IOV analysis yielded a DSC of 0.70 and HD95 of 30.14 mm. In clinical evaluation, DL-generated CTVs received a mean quality score of 3.38 out of 5, comparable to physician-delineated CTVs (3.13; P =0.13). Conclusion This study explored a novel application of DL in radiation oncology by developing algorithms for automated CTV segmentation in ISRT for high-risk pediatric HL. Clinical evaluation showed that the DL model was able to generate clinically useful CTVs with quality comparable to manually delineated CTVs, suggesting its potential to enhance contouring consistency and improve physician efficiency in ISRT planning. Publication History Article published online: 02 December 2025 © 2025. Thieme. All rights reserved. Georg Thieme Verlag KG Oswald-Hesse-Straße 50, 70469 Stuttgart, Germany
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".