Data-Efficient Lung Segmentation Using Foundational Models: Improving Clinical Workflow With Segment Anything Model (SAM) for Hyperpolarized Gas MRI
Bibliographic record
Abstract
Abstract Introduction: Hyperpolarized 129Xe/3 He lung MRI is a well-established method for evaluating pulmonary conditions and ventilation defects [1,2]. Traditional segmentation techniques [3] are effective but labor-intensive and time-consuming. Despite advances in deep learning (DL)-based segmentation of hyperpolarized gas MRI [4-6], the use of foundational models in this domain has been underutilized [7]. Our proof-of-concept research hypothesis is that foundational models, such as the Segment Anything Model (SAM) [8], can be utilized as efficient, ready-to-deploy for lung MRI segmentation in data-limited clinical environments. We propose that SAM can achieve superior segmentation performance compared to traditional convolutional neural networks (CNNs) even when trained on small datasets, making it a clinically viable tool for real-world use. Methods: We collected data from 56 participants (9 healthy, 28 COPD, 9 asthma, 10 COVID-19), resulting in 896 2D slices (128x128 pixels). The dataset was split 80% for training and 20% for testing. To assess data efficiency, experiments were performed with 25% of the training data. SAM was fine-tuned for segmentation using PyTorch on an NVIDIA GA102 GPU and compared to CNN-based models (UNet with ResNet18, ResNet152, VGG16, and VGG19). All models were evaluated on the same test set for consistency. Results: With only 25% of the available training data, SAM achieved a DSC of 0.97 for both proton and hyperpolarized gas MRI, significantly outperforming CNN models. UNet with VGG16 scored 0.89 for proton MRI and 0.78 for hyperpolarized gas MRI, while UNet with ResNet18 achieved 0.89 and 0.84, respectively. These results suggest that SAM can provide near-expert-level performance even in data-limited clinical settings. Conclusion: SAM outperformed traditional CNNs, particularly in data-limited conditions, demonstrating its potential as a robust, ready-to-deploy segmentation tool in clinical practice. Its ability to achieve high performance with minimal training data suggests SAM could be implemented in clinics where large datasets are not feasible, offering clinicians a fast, accurate, and scalable solution for lung MRI segmentation. Future work will explore one-shot learning techniques and larger datasets to enhance clinical applicability further. References: 1_Perron, S. et al. J_Magn_Reson_348, 107387 (2023).2_Kirby, M. et al. J_Appl_Physiol_114, 707-715 (2013). 3_Kirby, M. et al. Acad_Radiol_19, 141-152 (2012). 4_Astley, J. R. et al. Sci_Rep_12, 10566 (2022). 5_Astley, J. R. et al. J_Magn_Reson_Imaging_57, 1878-1890 (2022). 6_Astley, J. R. et al. Br_J_Radiol_95 (2022). 7_Babaeipour, R. et al. Bioengineering_10, 1349 (2023). 8_Kirillov, A. et al. ICCV, 3992-4003(2023).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".