Test-Retest Reliability of a New Medial Temporal Atrophy Morphological Metric
Bibliographic record
Abstract
Clinicians and researchers alike are in need of quantitative and robust measurement tools to assess medial temporal lobe atrophy (MTA) due to Alzheimer's disease (AD). We recently proposed a morphological metric, extracted from T1-weighted magnetic resonance images (MRI), to track and estimate MTA in cohorts of controls, AD, and mild cognitive impairment subjects, at high-risk of progression to dementia. In this paper, we investigated its reliability through analysis of within-session scan/repeat images and scan/rescans from large multicenter studies. In total, we used MRI data from 1051 subjects recruited at over 60 centers. We processed the data identically and calculated our metric for each individual, based on the concept of distance in a high-dimensional space of intensity and shape characteristics. Over 759 subjects, the scan/repeat change in the mean was 1.97% (SD: 21.2%). Over three subjects, the scan/rescan change in the mean was 0.89% (SD: 22.1%). At this level, the minimum trial size required to detect this difference is 68 individuals for both samples. Our scan/repeat and scan/rescan results demonstrate that our MTA assessment metric shows high reliability, a necessary component of validity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".