Systematic comparisons of different quality control approaches applied to three large pediatric neuroimaging datasets
Bibliographic record
Abstract
INTRODUCTION: Poor quality T1-weighted brain scans systematically affect the calculation of brain measures. Removing the influence of such scans requires identifying and excluding scans with noise and artefacts through a quality control (QC) procedure. While QC is critical for brain imaging analyses, it is not yet clear whether different QC approaches lead to the exclusion of the same participants. Further, the removal of poor-quality scans may unintentionally introduce a sampling bias by excluding the subset of participants who are younger and/or feature greater clinical impairment. This study had two aims: (1) examine whether different QC approaches applied to T1-weighted scans would exclude the same participants, and (2) examine how exclusion of poor-quality scans impacts specific demographic, clinical and brain measure characteristics between excluded and included participants in three large pediatric neuroimaging samples. METHODS: We used T1-weighted, resting-state fMRI, demographic and clinical data from the Province of Ontario Neurodevelopmental Disorders network (Aim 1: n = 553, Aim 2: n = 465), the Healthy Brain Network (Aim 1: n = 1051, Aim 2: n = 558), and the Philadelphia Neurodevelopmental Cohort (Aim 1: n = 1087; Aim 2: n = 619). Four different QC approaches were applied to T1-weighted MRI (visual QC, metric QC, automated QC, fMRI-derived QC). We used tetrachoric correlation and inter-rater reliability analyses to examine whether different QC approaches excluded the same participants. We examined differences in age, mental health symptoms, everyday/adaptive functioning, IQ and structural MRI-derived brain indices between participants that were included versus excluded following each QC approach. RESULTS: =0.52-0.59). Implementation of QC excluded younger participants, and tended to exclude those with lower IQ, and lower everyday/adaptive functioning scores across several approaches in a dataset-specific manner. Across nearly all datasets and QC approaches examined, excluded participants had lower estimates of cortical thickness and subcortical volume, but this effect did not differ by QC approach. CONCLUSION: The results of this study provide insight into the influence of QC decisions on structural pediatric imaging analyses. While different QC approaches exclude different subsets of participants, the variation of influence of different QC approaches on clinical and brain metrics is minimal in large datasets. Overall, implementation of QC tends to exclude participants who are younger, and those who have more cognitive and functional impairment. Given that automated QC is standardized and can reduce between-study differences, the results of this study support the potential to use automated QC for large pediatric neuroimaging datasets.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.109 | 0.244 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".