O2‐11‐02: Definition of harmonized protocol for hippocampal segmentation
Bibliographic record
Abstract
Heterogeneity of landmarks among protocols leads to different volume estimates, hampering comparison of studies and clinical use. There is an urgent need to define a harmonized protocol for manual hippocampal segmentation from magnetic resonance scans. Landmark differences among the 12 most common protocols were extracted, operationalized, and quantitatively investigated. The results were presented to the Delphi panel, consisting of seventeen researchers with substantial expertise in hippocampal segmentation, in order to reach an evidence-based consensus on segmentation landmarks. The Delphi panel participated in iterative anonymous voting sessions where feedback from previous rounds was utilized to progressively facilitate panelists' convergence on agreement. Panelists were presented with segmentation alternatives, each associated with quantitative data relating: (i) reliability, (ii) impact on whole hippocampal volume, and (iii) correlation with AD-related atrophy. Panelists were asked to choose among alternatives and provide justification, comments and level of agreement with the proposed solution. Anonymous votes and comments, and voting statistics of each round were fed into the following Delphi round. Exact probability on binomial tests of panelists' preferences was computed. Sixteen panelists completed four Delphi rounds. Agreement was significant on (i) inclusion of alveus/fimbria (P = 0.021); inclusion of the whole hippocampal tail (P = 0.013); (iii) segmentation of the medial border of the body following visible morphology as the first choice (P = 0.006) and following a horizontal line in the absence of morphological cues (P = 0.021); inclusion of the minimum hippocampus (comprising head and body) (P = 0.001); and inclusion of vestigial tissue in the segmentation of the tail (P = 0.022) (Figure). Significant agreement was also achieved for exclusion of internal cerebrospinal fluid pools (P = 0.004). Based on previous quantitative investigation, the hippocampus so defined covers 100% of hippocampal tissue, captures 100% of AD-related atrophy, and has good intra-rater (0.99) and inter-rater (0.94) reliability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".