Improving medical image segmentation with SAM2: analyzing the impact of object characteristics and finetuning on multi-planar datasets.
Bibliographic record
Abstract
This study investigates the factors that affect the performance of the Segment Anything Model 2 (SAM2) on medical imaging datasets, with a specific focus on the influence of object characteristics and the benefits of fine-tuning on multi-planar datasets. Utilizing data from three comprehensive medical imaging datasets—Medical Segmentation Decathlon (MSD), ISLES 2022, and BTCV Multi-Organ Abdominal Dataset—we analyzed SAM2's segmentation accuracy across a variety of object characteristics, such as size, intensity, location, and structural complexity. Our dataset included 714 cases, representing a various anatomical region and an independent test set of 985 video examples was used to validate our findings. Fine-tuning SAM2 led to notable improvements in segmentation performance across all metrics. The global mean Intersection over Union (IoU) increased from 0.690 to 0.827 while the Dice coefficient and Normalized Surface Dice (NSD) saw improvements of 15.58 % and 14.6 %, respectively. Challenging structures showed the most dramatic improvements, with the pancreas displaying a remarkable 48.8 % increase in Dice score and a 65.2 % improvement in IoU post-finetuning. Statistical analyses demonstrated significant correlations between segmentation performance and object characteristics. Medium-sized, centrally located structures with high solidity and smooth boundaries achieved the highest performance metrics. SAM2's segmentation performance is affected by object characteristics like size, location, and structural complexity. Fine-tuning the model with medical imaging data markedly enhances its accuracy, underlining SAM2's potential as a robust tool for clinical and research applications. The software, data, and resulting model are publicly accessible for non-commercial use. Our code will be released at: https://github.com/RadSam2/rad_sam2 • Object size, location, and structural complexity significantly affect SAM2's segmentation performance. • Fine-tuning SAM2 on medical imaging data leads to substantial improvements in Dice coefficient and Iou. • SAM2 demonstrates robust generalization and consistent performance across diverse anatomical regions post-fine-tuning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".