MétaCan
Menu
Back to cohort
Record W4385351255 · doi:10.1016/j.nicl.2023.103483

Validation of deep learning techniques for quality augmentation in diffusion MRI for clinical studies

2023· article· en· W4385351255 on OpenAlexafffund
Santiago Aja‐Fernández, Carmen Martín‐Martín, Álvaro Planchuelo‐Gómez, Abrar Faiyaz, Md Nasir Uddin, Giovanni Schifitto, Abhishek Tiwari, Saurabh J. Shigwan, Rajeev Kumar Singh, Tianshu Zheng, Zuozhen Cao, Dan Wu, Stefano B. Blumberg, Snigdha Sen, Tobias Goodwin-Allcock, Paddy J. Slator, Mehmet Yigit Avci, Zihan Li, Berkin Bilgiç, Qiyuan Tian, Xinyi Wang, Zihao Tang, Mariano Cabezas, Amelie Rauland, Dorit Merhof, Renata Manzano Maria, Vinícius P. Campos, Tales Santini, Marcelo A. C. Vieira, SeyyedKazem HashemizadehKolowri, Edward DiBella, Chenxu Peng, Zhimin Shen, Zan Chen, Irfan Ullah, Merry Mani, Hesam Abdolmotalleby, Samuel Eckstrom, Steven H. Baete, Patryk Filipiak, Tanxin Dong, Qiuyun Fan, Rodrigo de Luis‐García, Antonio Tristán‐Vega, Tomasz Pieciak

Bibliographic record

VenueNeuroImage Clinical · 2023
Typearticle
Languageen
FieldMedicine
TopicAdvanced Neuroimaging Techniques and Applications
Canadian institutionsWestern University
FundersScience and Technology Department of Zhejiang ProvinceNational Institute of Mental HealthAgencia Estatal de InvestigaciónEngineering and Physical Sciences Research CouncilNational Institute on AgingNarodowa Agencja Wymiany AkademickiejNational Institute of Neurological Disorders and StrokeNational Institute for Health and Care ResearchMinisterio de Ciencia e InnovaciónNational Institute of Biomedical Imaging and BioengineeringMinisterio de Ciencia, Innovación y UniversidadesMinistry of Science and Technology of the People's Republic of ChinaConselho Nacional de Desenvolvimento Científico e TecnológicoNational Natural Science Foundation of ChinaNational Institutes of HealthCanada Research ChairsUniversity College London Hospitals NHS Foundation TrustEuropean CommissionNational Institute of Dental and Craniofacial ResearchCoordenação de Aperfeiçoamento de Pessoal de Nível SuperiorDeutsche Forschungsgemeinschaft
KeywordsFalse positive paradoxChronic MigraineDiffusion MRIArtificial intelligenceGeneralizationMigraineWhite matterDeep learningMedicinePsychologyPattern recognition (psychology)Computer scienceStatisticsMagnetic resonance imagingMathematicsRadiologyInternal medicine

Abstract

fetched live from OpenAlex

The objective of this study is to evaluate the efficacy of deep learning (DL) techniques in improving the quality of diffusion MRI (dMRI) data in clinical applications. The study aims to determine whether the use of artificial intelligence (AI) methods in medical images may result in the loss of critical clinical information and/or the appearance of false information. To assess this, the focus was on the angular resolution of dMRI and a clinical trial was conducted on migraine, specifically between episodic and chronic migraine patients. The number of gradient directions had an impact on white matter analysis results, with statistically significant differences between groups being drastically reduced when using 21 gradient directions instead of the original 61. Fourteen teams from different institutions were tasked to use DL to enhance three diffusion metrics (FA, AD and MD) calculated from data acquired with 21 gradient directions and a b-value of 1000 s/mm2. The goal was to produce results that were comparable to those calculated from 61 gradient directions. The results were evaluated using both standard image quality metrics and Tract-Based Spatial Statistics (TBSS) to compare episodic and chronic migraine patients. The study results suggest that while most DL techniques improved the ability to detect statistical differences between groups, they also led to an increase in false positive. The results showed that there was a constant growth rate of false positives linearly proportional to the new true positives, which highlights the risk of generalization of AI-based tasks when assessing diverse clinical cohorts and training using data from a single group. The methods also showed divergent performance when replicating the original distribution of the data and some exhibited significant bias. In conclusion, extreme caution should be exercised when using AI methods for harmonization or synthesis in clinical studies when processing heterogeneous data in clinical studies, as important information may be altered, even when global metrics such as structural similarity or peak signal-to-noise ratio appear to suggest otherwise.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.007
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.323
Threshold uncertainty score0.869

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.007
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.459
GPT teacher head0.608
Teacher spread0.149 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations13
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueNeuroImage ClinicalSame topicAdvanced Neuroimaging Techniques and ApplicationsFrench-language works237,207