The development of a valid, reliable, harmonized segmentation protocol for hippocampal subfields and medial temporal lobe cortices: A progress update
Bibliographic record
Abstract
Abstract Background The medial temporal lobe (MTL, i.e. hippocampus and adjacent cortices) is particularly vulnerable to age‐related diseases: Alzheimer’s disease, other age‐related proteinothies (TDP‐43, AGD, etc) and vascular injury. Yet, the subregional pattern of vulnerability is thought to differ across etiologies; characterizing these differences using high‐resolution MRI may provide more insight in disease processes and better biomarkers. However, substantial differences in subfield definition has hindered the ability to compare results across laboratories or draw robust conclusions (Figure 1). The Hippocampal Subfields Group (HSG) is an international group seeking to remedy this problem by developing a histologically‐valid, reliable, and freely available segmentation protocol for high‐resolution T2‐weighted 3T MRI (http://www.hippocampalsubfields.com) Method Our workflow consists of five steps: 1) collecting histology samples labeled by multiple expert neuroanatomists to form a novel reference dataset to guide the development of the MRI segmentation protocol, 2) developing boundary definitions for each segment of the hippocampus, (head, body, and tail) and MTL cortices), 3) assessing HSG community agreement with boundary rules via online questionnaires and revising boundary rules based on questionnaire responses, and 4) testing reliability of the protocol definitions on multiple MRI datasets. Result For both the hippocampal body and head, we have developed a preliminary subfield segmentation protocol (i.e. completed steps 1‐2, see Figure 2 for a histology slice segmented by three anatomists). Step 3 was piloted for the outer boundaries of the body (i.e., the anterior/posterior, medial/lateral, and superior/inferior boundaries) using an online questionnaire describing each of the proposed rules. 29 labs participated and consensus agreement was reached for all rules, with only minor changes being made to improve comprehension and clarity. We are now creating and administering additional questionnaires for assessing agreement of the hippocampal body and head inner boundary rules (e.g., between the CA fields and dentate gyrus). Upon completion of the assessment/revision process for each set of rules, the final phase – reliability testing of the protocol – should begin mid 2020 for the body. Conclusion Once completed, the harmonized protocol will significantly facilitate cross‐study comparisons thus advancing insight in the role of hippocampal subfields across the lifespan in aging and disease.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.127 | 0.131 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.006 | 0.003 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.008 | 0.009 |
| Open science | 0.010 | 0.006 |
| Research integrity | 0.004 | 0.007 |
| Insufficient payload (model declined to judge) | 0.007 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".