Validation of an established deep learning auto-segmentation tool for cardiac substructures in 4D radiotherapy planning scans
Bibliographic record
Abstract
Background: Emerging data suggest that dose-sparing several key cardiac regions is prognostically beneficial in lung cancer radiotherapy. The cardiac substructures are challenging to contour due to their complex geometry, poor soft tissue definition on computed tomography (CT) and cardiorespiratory motion artefact. A neural network was previously trained to generate the cardiac substructures using three-dimensional radiotherapy planning CT scans (3D-CT). In this study, the performance of that tool on the average intensity projection from four-dimensional (4D) CT scans (4D-AVE), now commonly used in lung radiotherapy, was evaluated. Materials and Methods: The 4D-AVE of n=20 patients completing radiotherapy for lung cancer 2015-2020 underwent manual and automated cardiac substructure segmentation. Manual and automated substructures were compared geometrically and dosimetrically. Two senior clinicians also qualitatively assessed the auto-segmentation tool's output. Results: Geometric comparison of the automated and manual segmentations exhibited high levels of similarity across parameters, including volume difference (11.8% overall) and Dice similarity coefficient (0.85 overall), and were consistent with 3D-CT performance. Differences in mean (median 0.2 Gy, range -1.6-0.3 Gy) and maximum (median 0.4 Gy, range -2.2-0.9 Gy) doses to substructures were generally small. Nearly all structures (99.5 %) were deemed to be appropriate for clinical use without further editing. Conclusions: Cardiac substructure auto-segmentation using a deep learning-based tool trained on a 3D-CT dataset was feasible on the 4D-AVE scan, meaning this tool is suitable for use on 4D-CT radiotherapy planning scans. Application of this tool would increase the practicality of routine clinical cardiac substructure delineation, and enable further cardiac radiation effects research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".