Data-driven neuroanatomical subtypes of primary progressive aphasia
Bibliographic record
Abstract
The primary progressive aphasias are rare, language-led dementias, with three main variants: semantic, non-fluent/agrammatic and logopenic. Although the semantic variant has a clear neuroanatomical profile, the non-fluent/agrammatic and logopenic variants are difficult to discriminate from neuroimaging. Previous phenotype-driven studies have characterized neuroanatomical profiles of each variant on MRI. In this work, we used a machine learning algorithm known as SuStaIn to discover data-driven neuroanatomical 'subtype' progression profiles and performed an in-depth subtype-phenotype analysis to characterize the heterogeneity of primary progressive aphasia. Our study included 270 participants with primary progressive aphasia seen for research in the UCL Queen Square Institute of Neurology Dementia Research Centre, with follow-up scans available for 137 participants. This dataset included individuals diagnosed with all three main variants (semantic, n = 94; non-fluent/agrammatic, n = 109; logopenic, n = 51) and individuals with unspecified primary progressive aphasia (n = 16). A dataset of 66 patients (semantic, n = 37; non-fluent/agrammatic, n = 29) from the ARTFL LEFFTDS Longitudinal Frontotemporal Lobar Degeneration (ALLFTD) Research Study was used to validate our results. MRI scans were segmented, and SuStaIn was used on 19 regions of interest to identify neuroanatomical profiles independent of the diagnosis. We assessed the assignment of subtypes and stages, in addition to their longitudinal consistency. We discovered four neuroanatomical subtypes of primary progressive aphasia, labelled S1 (left temporal), S2 (insula), S3 (temporoparietal) and S4 (frontoparietal), exhibiting robustness to statistical scrutiny. S1 was correlated strongly with the semantic variant, whereas S2, S3 and S4 showed mixed associations with the logopenic and non-fluent/agrammatic variants. Notably, S3 displayed a neuroanatomical signature akin to a logopenic-only signature, yet a significant proportion of logopenic cases were allocated to S2. The non-fluent/agrammatic variant demonstrated diverse associations with S2, S3 and S4. No clear relationship emerged between any of the neuroanatomical subtypes and the unspecified cases. At first follow-up, subtype assignment was stable for 84% of patients, and stage assignment was stable for 91.9% of patients. We partially validated our findings in the ALLFTD dataset, finding comparable qualitative patterns. Our study, leveraging machine learning on a large primary progressive aphasia dataset, delineated four distinct neuroanatomical patterns. Our findings suggest that separable spatiotemporal neuroanatomical phenotypes do exist within the primary progressive aphasia spectrum, but that these are noisy, particularly for the non-fluent/agrammatic non-fluent/agrammatic and logopenic variants. Furthermore, these phenotypes do not always conform to standard formulations of clinico-anatomical correlation. Understanding the multifaceted profiles of the disease, encompassing neuroanatomical, molecular, clinical and cognitive dimensions, has potential implications for clinical decision support.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".