Multimodal subtypes identified in Alzheimer’s Disease Neuroimaging Initiative participants by missing-data-enabled subtype and stage inference
Bibliographic record
Abstract
Alzheimer's disease is a highly heterogeneous disease in which different biomarkers are dynamic over different windows of the decades-long pathophysiological processes, and potentially have distinct involvement in different subgroups. Subtype and Stage Inference is an unsupervised learning algorithm that disentangles the phenotypic heterogeneity and temporal progression of disease biomarkers, providing disease insight and quantitative estimates of individual subtype and stage. However, a key limitation of Subtype and Stage Inference is that it requires a complete set of biomarkers for each subject, reducing the number of datapoints available for model fitting and limiting applications of Subtype and Stage Inference to modalities that are widely collected, e.g. volumetric biomarkers derived from structural MRI. In this study, we adapted the Subtype and Stage Inference algorithm to handle missing data, enabling the application of Subtype and Stage Inference to multimodal data (magnetic resonance imaging, positron emission tomography, cerebrospinal fluid and cognitive tests) from 789 participants in the Alzheimer's Disease Neuroimaging Initiative. Missing-data Subtype and Stage Inference identified five subtypes having distinct progression patterns, which we describe by the earliest unique abnormality as 'Typical AD with Early Tau', 'Typical AD with Late Tau', 'Cortical', 'Cognitive' and 'Subcortical'. These new multimodal subtypes were differentially associated with age, years of education, Apolipoprotein E (APOE4) status, white matter hyperintensity burden and the rate of conversion from mild cognitive impairment to Alzheimer's disease, with the 'Cognitive' subtype showing the fastest clinical progression, and the 'Subcortical' subtype the slowest. Overall, we demonstrate that missing-data Subtype and Stage Inference reveals a finer landscape of Alzheimer's disease subtypes, each of which are associated with different risk factors. Missing-data Subtype and Stage Inference has broad utility, enabling the prediction of progression in a much wider set of individuals, rather than being restricted to those with complete data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".