Predicting Brain Amyloid using Multivariate Morphometry Statistics, Sparse Coding, and Correntropy: Validation in 1,101 Individuals from the ADNI and OASIS Databases
Bibliographic record
Abstract
ABSTRACT Biomarker-assisted preclinical/early detection and intervention in Alzheimer’s disease (AD) may be the key to therapeutic breakthroughs. One of the presymptomatic hallmarks of AD is the accumulation of beta-amyloid (Aβ) plaques in the human brain. However, current methods to detect Aβ pathology are either invasive (lumbar puncture) or quite costly and not widely available (amyloid PET). Our prior studies show that MRI-based hippocampal multivariate morphometry statistics (MMS) are an effective neurodegenerative biomarker for preclinical AD. Here we attempt to use MRI-MMS to make inferences regarding brain Aβ burden at the individual subject level. As MMS data has a larger dimension than the sample size, we propose a sparse coding algorithm, Patch Analysis-based Surface Correntropy-induced Sparse coding and max-pooling (PASCS-MP), to generate a low-dimensional representation of hippocampal morphometry for each subject. Then we apply these individual representations and a binary random forest classifier to predict brain Aβ positivity for each person. We test our method in two independent cohorts, 841 subjects from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) and 260 subjects from the Open Access Series of Imaging Studies (OASIS). Experimental results suggest that our proposed PASCS-MP method and MMS can discriminate Aβ positivity in people with mild cognitive impairment (MCI) (Accuracy (ACC)=0.89 (ADNI)) and in cognitively unimpaired (CU) individuals (ACC=0.79 (ADNI) and ACC=0.81 (OASIS)). These results compare favorably relative to measures derived from traditional algorithms, including hippocampal volume and surface area, shape measures based on spherical harmonics (SPHARM), and our prior Patch Analysis-based Surface Sparse-coding and Max-Pooling (PASS-MP) methods.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".