UA-UNet: Uncertainty Aware Pseudo-Label Generation in Residual U-Net for Medical Image Segmentation
Bibliographic record
Abstract
Medical image segmentation is crucial for accurate diagnosis and treatment planning. Most segmentation models are developed through supervised learning, which requires access to large amounts of annotated data. However, acquiring such datasets is often expensive and time-consuming. Semi-supervised learning (SSL) approaches alleviate this challenge by leveraging both labeled and unlabeled data to improve model performance. While SSL methods have achieved promising results, they still have some limitations. Predictions made by these methods disregard the model uncertainty, leading to unreliable predictions. The predictions made by these models are often affected by the noise introduced during the SSL process. Additionally, these models fail to adapt well to complex medical images and are unable to scale well in out-of-domain samples. To address these limitations, we propose UA-UNet, an uncertainty-guided teacher-student model architecture built on top of the Residual U-Net for multi-class medical image segmentation. The model incorporates uncertainty estimation, which guides the generation of high-quality pseudo-labels from an ensemble of teacher models. By combining consistency regularization and pseudo-labeling, our method effectively reduces the influence of high-uncertainty regions while enhancing segmentation accuracy. We compared the model with 10 other methods on the 2023 Kidney Tumor Segmentation Challenge dataset (KiTS23). The proposed approach outperformed state-of-the-art models with a Dice score of 0.901 and IoU of 0.891. The proposed model also provides uncertainty maps, which could enhance the interpretability of the segmentation result. These features make UA-UNet a robust method for semi-supervised segmentation in medical imaging, particularly when labeled data is scarce.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".