Do we really need to normalize Hippocampal volume?
Bibliographic record
Abstract
Abstract Background Hippocampal volume (HCvol) is an important biomarker in the study of neurodegeneration (1‐3). A drawback of HCvol as a clinical biomarker is its high variability across the population (4). While the common approach to account for this variability is to normalize for head‐size variability using the intracranial volume (ICV) (5), other approaches use the idea of ex‐vacuo dilation, considering the expansion of the surrounding ventricle (6,7). Method With a library of 80 manually segmented hippocampi and surrounding temporal horns of the lateral ventricles from healthy subjects we trained three automatic segmentation methods: multi‐atlas label fusion (MALF) (8), non‐local patch‐based segmentation (NLPB, aka SNIPE) (9) and a Convolutional Neural Network (CNN) with a U‐Net architecture (10) and compared their performance (Cohen’s Kappa). We obtained the raw and normalized Hippocampal volumes (HCvol/ICV) as well as the Hippocampal‐to‐Ventricle Ratio (HVR = HCvol/(HCvol+CSFvol)) (7) on the baseline T1w MRI scans from the ADNI dataset for each segmentation method. We used these measures to calculate the robust effect sizes (Cohen’s d) and their confidence intervals (bootstrap: 5,000 resamples) between cognitively normal (CN), mild cognitive impaired (MCI) and subjects with dementia (AD) from ADNI‐1, 2, 3 and Go. Result After acquisition QC (CN = 502, MCI = 815, AD = 324), Table 1 shows the number of successful HC and CSF segmentations for each method. The CNN U‐Net had the least number of failed segmentations. The CNN U‐Net also obtained the highest performance. While NLPB had similar Kappa values to MALF, it had less variability (Fig 1). On Fig 2, we show the calculated effect sizes between Diagnoses, HC measures and segmentation methods. HVR yields a greater effect size than HCvol and HCvol/ICV to separate AD from CN or MCI when using MALF. With NLPB, HVR and HCvol are better than HCvol/ICV. Finally, with the CNN U‐NET, the three metrics perform similarly. Conclusion The CNN U‐Net provides improved segmentations over both MALF and NLPB. Such better‐quality segmentations reduce the need for normalization. However, HVR might be better suited than ICV normalized values when working with limited quality segmentations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.088 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".