MétaCan
Menu
Back to cohort
Record W1979051296 · doi:10.3389/fninf.2015.00012

Reproducibility of neuroimaging analyses across operating systems

2015· article· en· W1979051296 on OpenAlexafffund
Tristan Glatard, Lindsay B. Lewis, Rafael Ferreira da Silva, Reza Adalat, Natacha Beck, Claude Lepage, Pierre Rioux, Marc-Étienne Rousseau, Tarek Sherif, Ewa Deelman, Najmeh Khalili‐Mahani, Alan C. Evans

Bibliographic record

VenueFrontiers in Neuroinformatics · 2015
Typearticle
Languageen
FieldNeuroscience
TopicFunctional Brain Connectivity Studies
Canadian institutionsMontreal Neurological Institute and HospitalMcGill University
FundersNational Institute on AgingNational Institute of Biomedical Imaging and BioengineeringCanadian Institutes of Health ResearchUniversity of California, San DiegoGenentechNational Institutes of HealthIXICOServierNovartis Pharmaceuticals CorporationLabEx PRIMESEisaiNorthern California Institute for Research and EducationF. Hoffmann-La RocheAgence Nationale de la RechercheSynarcUniversity of Southern CaliforniaMedpaceUniversité de LyonBiogenPfizerCompute CanadaU.S. Department of DefenseEli Lilly and CompanyBristol-Myers SquibbFoundation for the National Institutes of HealthAlzheimer's Disease Neuroimaging InitiativeMeso Scale DiagnosticsAlzheimer's Drug Discovery FoundationNational Science Foundation
KeywordsComputer scienceReproducibilityNeuroimagingPipeline transportArtificial intelligencePattern recognition (psychology)NeuroscienceMathematicsPsychologyStatistics

Abstract

fetched live from OpenAlex

Neuroimaging pipelines are known to generate different results depending on the computing platform where they are compiled and executed. We quantify these differences for brain tissue classification, fMRI analysis, and cortical thickness (CT) extraction, using three of the main neuroimaging packages (FSL, Freesurfer and CIVET) and different versions of GNU/Linux. We also identify some causes of these differences using library and system call interception. We find that these packages use mathematical functions based on single-precision floating-point arithmetic whose implementations in operating systems continue to evolve. While these differences have little or no impact on simple analysis pipelines such as brain extraction and cortical tissue classification, their accumulation creates important differences in longer pipelines such as subcortical tissue classification, fMRI analysis, and cortical thickness extraction. With FSL, most Dice coefficients between subcortical classifications obtained on different operating systems remain above 0.9, but values as low as 0.59 are observed. Independent component analyses (ICA) of fMRI data differ between operating systems in one third of the tested subjects, due to differences in motion correction. With Freesurfer and CIVET, in some brain regions we find an effect of build or operating system on cortical thickness. A first step to correct these reproducibility issues would be to use more precise representations of floating-point numbers in the critical sections of the pipelines. The numerical stability of pipelines should also be reviewed.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.027
metaresearch head score (Gemma)0.142
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.973
Threshold uncertainty score0.143

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0270.142
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.002
Scholarly communication0.0030.003
Open science0.0020.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0040.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.146
GPT teacher head0.357
Teacher spread0.211 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations149
Published2015
Admission routes2
Has abstractyes

Explore more

Same venueFrontiers in NeuroinformaticsSame topicFunctional Brain Connectivity StudiesFrench-language works237,207