Generalizability of deep learning in organ-at-risk segmentation: A transfer learning study in cervical brachytherapy
Bibliographic record
Abstract
PURPOSE: Deep learning can automate delineation in radiation therapy, reducing time and variability. Yet, its efficacy varies across different institutions, scanners, or settings, emphasizing the need for adaptable and robust models in clinical environments. Our study demonstrates the effectiveness of the transfer learning (TL) approach in enhancing the generalizability of deep learning models for auto-segmentation of organs-at-risk (OARs) in cervical brachytherapy. METHODS: A pre-trained model was developed using 120 scans with ring and tandem applicator on a 3T magnetic resonance (MR) scanner (RT3). Four OARs were segmented and evaluated. Segmentation performance was evaluated by Volumetric Dice Similarity Coefficient (vDSC), 95 % Hausdorff Distance (HD95), surface DSC, and Added Path Length (APL). The model was fine-tuned on three out-of-distribution target groups. Pre- and post-TL outcomes, and influence of number of fine-tuning scans, were compared. A model trained with one group (Single) and a model trained with all four groups (Mixed) were evaluated on both seen and unseen data distributions. RESULTS: TL enhanced segmentation accuracy across target groups, matching the pre-trained model's performance. The first five fine-tuning scans led to the most noticeable improvements, with performance plateauing with more data. TL outperformed training-from-scratch given the same training data. The Mixed model performed similarly to the Single model on RT3 scans but demonstrated superior performance on unseen data. CONCLUSIONS: TL can improve a model's generalizability for OAR segmentation in MR-guided cervical brachytherapy, requiring less fine-tuning data and reduced training time. These results provide a foundation for developing adaptable models to accommodate clinical settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".