Vasculature segmentation in 3D hierarchical phase-contrast tomography images of human kidneys
Bibliographic record
Abstract
Efficient algorithms are needed to segment vasculature in new 3D medical imaging datasets at scale for research and clinical applications. Manual segmentation of vessels in images is time-consuming and expensive whereas computational approaches have limited accuracy. We organize a global machine learning competition, engaging 1,401 participants, to promote development of deep learning methods for 3D blood vessel segmentation in Hierarchical Phase-Contrast Tomography (HiP-CT) datasets. This paper presents a meta-analysis of the top-performing solutions, focusing on segmentation accuracy and morphological analysis. The competition and subsequent analysis reveal convergent methodological innovations: pseudo-labeling approaches that exploit data distributions, metrics and loss functions that optimize for vessel surface and topology, and multi-scale approaches that handle data heterogeneity. Additionally, the paper presents techniques for building deep learning models for the defined task, metrics to assess and compare algorithm performance, and a dataset with manually annotated and curated gold standard segmentations for future studies in blood vessel segmentation within HiP-CT imaging.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".