Multi-modality Artificial Intelligence for Involved-Site Radiation Therapy: Clinical Target Volume Delineation in High-Risk Pediatric Hodgkin Lymphoma
Notice bibliographique
Résumé
Introduction Clinical target volume (CTV) delineation for involved-site radiation therapy (ISRT) in Hodgkin lymphoma (HL) is time-consuming due to the need to analyze multi-time-point PET/CT scans co-registered to the planning CT. Deep learning (DL) has the potential to streamline this task, but its feasibility remains unexplored. Our goal was to develop automated CTV segmentation algorithms that integrated multi-modality imaging to facilitate ISRT planning. Methods This study included planning CT, baseline PET/CT (PET1), and interim PET/CT (PET2) scans from 288 pediatric patients with high-risk HL in the Children’s Oncology Group AHOD 1331 trial. Data from 58 patients across 24 institutions were held out for external testing, while the remaining 230 cases from 95 institutions were used for model development. We investigated three DL architectures (SegResNet, ResUNet, and SwinUNETR) and evaluated the impact of incorporating PET1 and PET2 images alongside the planning CT. Performance was assessed using the 95th percentile Hausdorff distance (HD95) and Dice similarity coefficient (DSC). Inter-observer variability (IOV) was estimated by comparing original institutional CTVs with those newly delineated by four board-certified radiation oncologists on a subset of 10 cases. The quality of CTVs generated by the top-performing model and those from original institutions was independently assessed on 40 other cases by four radiation oncologists, who were blinded to the source of the CTVs. Results On the external cohort, a SwinUNETR model incorporating planning CT, PET1, and PET2 images achieved the highest performance, with an HD95 of 34.43 mm, and DSC of 0.72. In comparison, the best planning CT-only model attained an HD95 of 58.94 mm and DSC of 0.68. All models incorporating PET/CT images were significantly better (P<0.01) than CT-only models. IOV analysis yielded a DSC of 0.70 and HD95 of 30.14 mm. In clinical evaluation, DL-generated CTVs received a mean quality score of 3.38 out of 5, comparable to physician-delineated CTVs (3.13; P =0.13). Conclusion This study explored a novel application of DL in radiation oncology by developing algorithms for automated CTV segmentation in ISRT for high-risk pediatric HL. Clinical evaluation showed that the DL model was able to generate clinically useful CTVs with quality comparable to manually delineated CTVs, suggesting its potential to enhance contouring consistency and improve physician efficiency in ISRT planning. Publication History Article published online: 02 December 2025 © 2025. Thieme. All rights reserved. Georg Thieme Verlag KG Oswald-Hesse-Straße 50, 70469 Stuttgart, Germany
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,003 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».