MétaCan
Menu
Back to cohort
Record W3004310692 · doi:10.1101/2020.01.28.922971

Reliability Assessment of Tissue Classification Algorithms for Multi-Center and Multi-Scanner Data

2020· preprint· en· W3004310692 on OpenAlexafffund
Mahsa Dadar, Simon Duchesne

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2020
Typepreprint
Languageen
FieldComputer Science
TopicMedical Image Segmentation Techniques
Canadian institutionsUniversité Laval
FundersConsortium canadien en neurodégénérescence associée au vieillissement
KeywordsScannerSegmentationRobustness (evolution)Artificial intelligenceSample size determinationComputer sciencePattern recognition (psychology)Brain sizeStatisticsMathematicsNuclear medicineMedicineMagnetic resonance imagingBiologyRadiology

Abstract

fetched live from OpenAlex

Abstract Background Gray and white matter volume difference and change are important imaging markers of pathology and disease progression in neurology and psychiatry. Such measures are usually estimated from tissue segmentation maps produced by publicly available image processing pipelines. However, the reliability of the produced segmentations when using multi-center and multi-scanner data remains understudied. Here, we assess the robustness of six publicly available tissue classification pipelines across images acquired from different MR scanners and sites. Methods We used T1-weighted images of a single individual, scanned in 73 sessions across 27 different sites to assess the robustness of the tissue classification tools. Variability in Dice Kappa values and tissue volumes was assessed for Atropos, BISON, Classify_Clean, FAST, FreeSurfer, and SPM12. We also estimated the sample size necessary to detect a significant 1% volume reduction based on the variability of the estimates from each method within and across scanner models. Results BISON had the lowest overall variability in its volumetric estimates, followed by FreeSurfer, and SPM12. All methods also had significant differences between some of their estimates across different scanner manufacturers (e.g. BISON had significantly higher GM estimates and correspondingly lower WM estimates for GE scans compared to Philips and SIEMENs), and different signal-to-noise ratio (SNR) levels (e.g. FAST and FreeSurfer had significantly higher WM volume estimates for high versus medium and low SNR tertiles as well as correspondingly lower GM volume estimates). BISON also had the smallest sample size requirement across all scanners and tissue types, followed by FreeSurfer, and SPM12. Conclusions Our comparisons provide a benchmark on the reliability of the publicly used tissue classification techniques and the amount of variability that can be expected when using large multi-center datasets and multi-scanner databases.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.083
metaresearch head score (Gemma)0.182
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.083
Threshold uncertainty score0.440

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0830.182
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0040.003
Science and technology studies0.0020.002
Scholarly communication0.0030.002
Open science0.0030.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.113
GPT teacher head0.363
Teacher spread0.250 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2020
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicMedical Image Segmentation TechniquesFrench-language works237,207