Impact of target volume segmentation accuracy and variability on treatment planning for 4D-CT-based non-small cell lung cancer radiotherapy
Bibliographic record
Abstract
BACKGROUND: Accurate target volume segmentation is crucial for success in image-guided radiotherapy. However, variability in anatomical segmentation is one of the most significant contributors to uncertainty in radiotherapy treatment planning. This is especially true for lung cancer where target volumes are subject to varying magnitudes of respiratory motion. MATERIAL AND METHODS: This study aims to analyze multiple observer target volume segmentations and subsequent intensity-modulated radiotherapy (IMRT) treatment plans defined by those segmentations against a reference standard for lung cancer patients imaged with four-dimensional computed tomography (4D-CT). Target volume segmentations of 10 patients were performed manually by six physicians, allowing for the calculation of ground truth estimate segmentations via the simultaneous truth and performance level estimation (STAPLE) algorithm. Segmentation variability was assessed in terms of distance- and volume-based metrics. Treatment plans defined by these segmentations were then subject to dosimetric evaluation consisting of both physical and radiobiological analysis of optimized 3D dose distributions. RESULTS: Significant differences were noticed amongst observers in comparison to STAPLE segmentations and this variability directly extended into the treatment planning stages in the context of all dosimetric parameters used in this study. Mean primary tumor control probability (TCP) ranged from (22.6±11.9)% to (33.7±0.6)%, with standard deviation ranging from 0.5% to 11.9%. However, mean normal tissue complication probabilities (NTCP) based on treatment plans for each physician-derived target volume well as the NTCP derived from STAPLE-based treatment plans demonstrated no discernible trends and variability appeared to be patient-specific. This type of variability demonstrated the large-scale impact that target volume segmentation uncertainty can play in IMRT treatment planning. CONCLUSIONS: Significant target volume segmentation and dosimetric variability exists in IMRT treatment planning amongst experts in the presence of a reference standard for 4D-CT-based lung cancer radiotherapy. Future work is needed to mitigate this uncertainty and ensure highly accurate and effective radiotherapy for lung cancer patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".