MétaCan
Menu
← Back to cohort
Record W7084098394 · doi:10.6084/m9.figshare.30059055

Additional file 1 of Lung lobe segmentation: performance of open-source MOOSE, TotalSegmentator, and LungMask models compared to a local in-house model

2025· article· en· W7084098394 on OpenAlexaff

Bibliographic record

VenueOpen MIND · 2025
Typearticle
Languageen
FieldImmunology and Microbiology
TopicT-cell and B-cell Immunology
Canadian institutionsOttawa HospitalUniversity of OttawaCarleton University
Fundersnot available
KeywordsSegmentationBonferroni correctionMissing dataIntersection (aeronautics)Pattern recognition (psychology)Data set

Abstract

fetched live from OpenAlex

Additional file 1: Table S1. Criteria used to assign a segmentation difficulty category to each chest CT scan. Fig. S1. Flowchart illustrating the data collection process, including the total number of cases initially collected, the exclusion criteria applied, and the final dataset used for analysis. Fig. S2. Boxplots displaying the performance metrics for open-source models across different difficulty categories for all lobes combined. From left to right, the box plots represent the performance evaluated using the DSC, rHd95, and NSD metrics, respectively. All p-values are after a Bonferroni correction factor of 3 was applied (n = 129 images; 42 Easy, 41 moderate, and 46 hard, ×5 segments per image). Fig. S3. Comparative analysis of segmentation accuracy of hard cases from the internal test set (n = 16), specifically those without missing lobes (n = 9) and those with missing lobes (n = 7), indicating that missing lobes have a disproportionately large impact on segmentation accuracy for all models except our model. Fig. S4. Bar chart comparing IoU scores for five lung lobes across 55 cases from the LOLA11 dataset. The chart is divided into two panels: Cases 1−28 and Cases 29−55. Each bar represents the IoU score for a specific lobe (LUL, LLL, RUL, RML, RLL) in each case, with scores ranging from -1.0 to 1.0, where -1.0 indicates excluded evaluations, 0 indicates no overlap, and 1.0 indicates perfect overlap between the model prediction and the ground truth. IoU, Intersection over union; LLL, Left Lower Lobe; LUL, Left Upper Lobe; RML, Right Middle Lobe; RLL, Right Lower Lobe; RUL, Right Upper Lobe. Fig. S5. Qualitative results showcasing some of the instances of the LOLA11 challenge dataset where our model successfully performed. Conditions included noisy scans, low-resolution images, emphysematous changes, and lesions ranging from small to moderate in size, both cystic and solid. Red arrows indicate areas where the model failed to make accurate predictions. LLL, Left lower lobe; LOLA11, LObe and Lung Analysis 2011; LUL, Left upper lobe; RLL, Right lower lobe; RML, Right middle lobe; RUL, Right upper lobe. Fig. S6. Qualitative results showcasing instances where our model struggled to perform accurately. The conditions and pathologies presented in these cases include disease- and treatment-related missing lobes (cases 44, 45, and 48), disease-related volume loss (cases 06, 20, 31, and 52), large lesions (case 31), and emphysematous changes (case 28). Red arrows indicate areas where the model failed to make accurate predictions, including misidentified fissures, exclusion of high-density parenchymal and pleural regions, and false positive class predictions. LLL, Left lower lobe; LUL, Left upper lobe; RLL, Right lower lobe; RML, Right middle lobe; RUL, Right upper lobe. Fig. S7. Two cases from the LOLA11 challenge where no overlap between our model’s prediction and ground truth annotations was recorded, with an IoU of 0 reported for the RML in case 21, and the LLL in case 52. In case 52, the red arrow highlights the ground truth region labeled as LLL, while the blue arrow indicates the location of the left oblique fissure. This visual interpretation calls into doubt the accuracy of ground truth segmentation and/or annotations in the LOLA11 competition. IoU, Intersection over union; LLL, Left lower lobe; LOLA11, LObe and Lung Analysis 2011; LUL, Left upper lobe; RLL, Right lower lobe; RML, Right middle lobe; RUL, Right upper lobe. Fig. S8. Example slices of LOLA11 challenge cases with missing right lung (case 44), missing left lung (case 45), and missing RML (case 48), demonstrating varying degrees of segmentation accuracy by the four models evaluated. The red arrow indicates a structure resembling the right horizontal fissure. LLL, Left lower lobe; LOLA11, LObe and Lung Analysis 2011; LUL, Left upper lobe; RLL, Right lower lobe; RML, Right middle lobe; RUL, Right upper lobe.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.025
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.601
Threshold uncertainty score0.569

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0030.025
Meta-epidemiology (narrow)0.0030.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.000
Scholarly communication0.0020.002
Open science0.0030.002
Research integrity0.0020.001
Insufficient payload (model declined to judge)0.6010.163

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.019
GPT teacher head0.267
Teacher spread0.247 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueOpen MIND→Same topicT-cell and B-cell Immunology→French-language works237,207→