Automatic deep learning segmentation of the hippocampus on high‐resolution diffusion magnetic resonance imaging and its application to the healthy lifespan
Bibliographic record
Abstract
Abstract Diffusion tensor imaging (DTI) can provide unique contrast and insight into microstructural changes with age or disease of the hippocampus, although it is difficult to measure the hippocampus because of its comparatively small size, location, and shape. This has been markedly improved by the advent of a clinically feasible 1‐mm isotropic resolution 6‐min DTI protocol at 3 T of the hippocampus with limited brain coverage of 20 axial‐oblique slices aligned along its long axis. However, manual segmentation is too laborious for large population studies, and it cannot be automatically segmented directly on the diffusion images using traditional T 1 or T 2 image‐based methods because of the limited brain coverage and different contrast. An automatic method is proposed here that segments the hippocampus directly on high‐resolution diffusion images based on an extension of well‐known deep learning architectures like UNet and UNet++ by including additional dense residual connections. The method was trained on 100 healthy participants with previously performed manual segmentation on the 1‐mm DTI, then evaluated on typical healthy participants (n = 53), yielding an excellent voxel overlap with a Dice score of ~ 0.90 with manual segmentation; notably, this was comparable with the inter‐rater reliability of manually delineating the hippocampus on diffusion magnetic resonance imaging (MRI) (Dice score of 0.86). This method also generalized to a different DTI protocol with 36% fewer acquisitions. It was further validated by showing similar age trajectories of volumes, fractional anisotropy, and mean diffusivity from manual segmentations in one cohort (n = 153, age 5–74 years) with automatic segmentations from a second cohort without manual segmentations (n = 354, age 5–90 years). Automated high‐resolution diffusion MRI segmentation of the hippocampus will facilitate large cohort analyses and, in future research, needs to be evaluated on patient groups.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".