Assessing the Performance of Foundation Models in Prostate Segmentation Across Different Ultrasound Modalities
Bibliographic record
Abstract
Prostate cancer (PCa) remains a major health issue for men and relies on early diagnostic procedures such as transrectal ultrasound-guided (TRUS) biopsy for effective treatment. Since this procedure is prone to high rates of false negatives, advancements are being made in ultrasound (US) imaging and artificial intelligence to improve the detection of cancerous lesions. Segmentation is necessary for establishing the area for analysis however, manual segmentation can be laborious and time-intensive. Recent studies have focused on leveraging pre-trained foundation models that can be applied to various downstream tasks such as prostate segmentation in US imaging. However, only the processed B-mode US images were used in these studies and other imaging modes like raw RF images, although proved superior in tissue characterization, have not been explored. This study evaluates the segmentation capabilities of SAM-Med2D, a foundation model finetuned on a large-scale medical dataset using different ultrasound modalities. We finetune the pre-trained model (baseline) with and without adapter layers on prostate US data. Furthermore, we assess the results across different anatomical locations of the prostate. Our findings show that B-mode base data are more compatible with the SAM-Med2D model. Also, we observe a substantial improvement in model performance over baseline when finetuning with a small dataset, slightly better performance when employing adapter layers, and no substantial difference between anatomical locations. Overall, our work highlights the potential application of SAM-Med2D in various US modalities, which is essential for RF-based pipelines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".