MétaCan
Menu
Back to cohort
Record W4416438657 · doi:10.3174/ajnr.a9110

Hippocampal Segmentation Performance on 7T MRI: Intensity-Based Accuracy Assessment with Paired 3T–7T Volume Comparison across Multiple Algorithms

2025· article· en· W4416438657 on OpenAlexaboutno aff
Justin Cramer, Ichiro Ikuta, Leslie C. Baxter, Jonathon J. Parker, Yalin Wang, Yuxiang Zhou

Bibliographic record

VenueAmerican Journal of Neuroradiology · 2025
Typearticle
Languageen
FieldComputer Science
TopicMedical Image Segmentation Techniques
Canadian institutionsnot available
Fundersnot available
KeywordsSegmentationHippocampal formationPattern recognition (psychology)Volume (thermodynamics)

Abstract

fetched live from OpenAlex

ABSTRACT BACKGROUND AND PURPOSE: Clinical adoption of 7 Tesla (7T) MRI is increasing, yet the performance of commonly used hippocampal segmentation algorithms, none of which are trained on 7T data, remains largely uncharacterized. This study evaluates segmentation accuracy at 7T using a voxel intensity–based method and examines volumetric differences between paired 3T and 7T hippocampal segmentations. MATERIALS AND METHODS: Two retrospective datasets from a single center were analyzed. For the 7T-only accuracy assessment cohort, 269 brain MRI studies performed on a Siemens Magnetom Terra.X with paired pre-and post-contrast T1 MPRAGE sequences (0.6 mm isovoxel) were utilized. For the 3T–7T cross-field comparison cohort, 39 unique subjects were identified with both 3T and 7T precontrast T1 MPRAGE sequences. Hippocampal segmentation was performed on the 7T-only cohort with AssemblyNet, e2dhipseg, FastSurfer, HippMapper, hippodeep, and QuickNat. QuickNat was removed from the 3T-7T cohort due to poor performance at 7T, and NeuroQuant 5.0 was added. Voxel intensity–based correction metrics quantified segmentation accuracy at 7T, with lower total correction volumes indicating better performance. Paired 3T–7T volume differences were assessed using the Wilcoxon signed-rank test, and corresponding NeuroQuant normative percentiles were also compared. RESULTS: At 7T, hippodeep achieved the lowest total correction volume (0.58 mL), followed by e2dhipseg (0.67 mL), FastSurfer (0.78 mL), HippMapper (0.84 mL), and AssemblyNet (0.89 mL). Welch ANOVA with Tukey post-hoc testing confirmed significant pairwise differences between all algorithms (p < 0.001). In the paired 3T–7T analysis, all algorithms yielded significantly smaller 7T volumes (p < 0.001), with mean absolute differences ranging from 0.19 mL (hippodeep) to 1.54 mL (HippMapper). NeuroQuant volumes differed by 0.54 mL, corresponding to a mean 41-point shift in normative percentiles. CONCLUSIONS: Hippodeep required the least total correction at 7T and had the smallest 3T–7T volume differences, suggesting it offers the most consistent cross-field performance among tested methods. However, consistent 7T volumetric underestimation across algorithms and the associated large normative percentile shifts from small volume changes indicate that a dedicated 7T normative database is necessary for meaningful clinical use. ABBREVIATIONS: MNI = Montreal Neurological Institute; HSD = honestly significant difference; T = Tesla.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.009
metaresearch head score (Gemma)0.020
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.009
Threshold uncertainty score0.046

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0090.020
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0020.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.336
Teacher spread0.318 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueAmerican Journal of NeuroradiologySame topicMedical Image Segmentation TechniquesFrench-language works237,207