Quantifying brain atrophy in Frontotemporal Dementia: a head-to-head comparison of neuroimaging techniques*
Bibliographic record
Abstract
Frontotemporal Dementia (FTD) is a neurodegenerative disorder characterized by extensive atrophy in the frontal and temporal lobes of the brain as well as high cerebrovascular burden. While anatomical Magnetic Resonance Imaging (MRI) is well established for quantifying brain atrophy in FTD, the variability in (pre-)processing methods limit the generalizability and comparability of findings. This study systematically compared the robustness and sensitivity of multiple widely used neuroimaging metrics, namely Deformation-Based Morphometry (DBM), Voxel-Based Morphometry (VBM), Cortical Thickness (CT), and segmentation-based grey matter Volumes, in detecting atrophy across FTD subtypes. We processed 732 T1-weighted MRI scans from 156 participants with FTD and 139 healthy controls from the Frontotemporal Lobar Degeneration Neuroimaging Initiative using our in-house pipeline PELICAN (Dadar et al., 2025) for volumetric measures and Free Surfer version 7 (Fischl, 2012) for CT and grey matter segmentations. Visual quality control using consistent quality control images at each step of the pipelines revealed significantly higher failure rates for CT (38.52%) and Free Surfer segmentations (23.63%) relative to PELICAN’s volumetric measures (2.04% DBM, 3.05% VBM). Failure rates differed between FTD subtypes and were related to pathological burden. Particularly for Free Surfer, errors occurred predominantly in regions with high prevalence of atrophy and White Matter Hyperintensities. In PELICAN, the addition of a FTD-specific template as an intermediate step during nonlinear registration decreased the failure rates in this step in the FTD population. We then applied linear regression models to assess each metric’s sensitivity in detecting cross-sectional differences between FTD groups controls as well as linear mixed-effects models to determine which method is most sensitive to longitudinal anatomical changes. While CT yielded effect sizes comparable to VBM and DBM when analyzing the same subset of successfully processed scans, VBM and DBM demonstrated enhanced power to detect effects due to lower failure rates and higher participant retention in the full sample. Overall, we demonstrate that image processing methodology and pipeline selection profoundly influences effect sizes and statistical power to detect meaningful between-group differences or longitudinal changes. Volumetric measures (DBM and VBM) yielded sufficiently robust pipeline outcomes to maintain adequate statistical power for capturing atrophy patterns after quality control procedures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".