Validation and replication of a prognostic machine learning model for enrichment of cognitive decliners in clinical trials
Bibliographic record
Abstract
Abstract Background Identifying individuals who will experience cognitive decline is paramount to evaluating novel treatments in Alzheimer’s clinical trials. We propose to use highly specific neuroimaging and genetic signatures that are indicative of the risk of cognitive decline in individuals with mild Alzheimer’s disease and mild cognitive impairment. Method Baseline measurements of gray matter volumes, from structural magnetic resonance imaging, APOE4 status, age and sex were provided to a machine learning prognostic enrichment solution Cog‐Foresight from Perceiv Research Inc. Canada to identify signatures of individuals at risk of cognitive decline up to two years of follow‐up. The resulting three cohorts of individuals were: (a) low‐risk, (b) moderate‐risk, (c) high‐risk of decline according to the enrichment tool. ADAS‐cog11, CDR‐SB and MMSE were used as clinical endpoints to evaluate the trajectory of each cohort. The prognostic tool was evaluated on three independent datasets: ADNI1 (n=345), ADNI2 (n=238) [1], and the placebo arm of an 18‐month Phase 3 study in mild AD (n=110). Result From the 693 individuals selected from the three datasets, 40% were deemed low‐risk and 23% were deemed high‐risk. In ADNI1, the mean rate of change at 18 months from baseline for the CDR‐SB for the low‐, moderate‐, and high‐risk groups were 0.51 points, 1.08 points, and 1.83 points, respectively. In ADNI2, the mean rate of change at 18 months from baseline for the CDR‐SB for the low‐, moderate‐, and high‐risk groups were 0.09 points, 0.57 points, and 1.79 points, respectively. In the Phase 3 study, the mean rate of change at 18 months for the CDR‐SB for the low‐, moderate‐, and high‐risk groups were 0.5 points, 0.88 points, and 2.04 points, respectively. Similar trajectories and conclusions were found for all other endpoints. Conclusion The use of multimodal biomarkers with a machine learning‐based targeted selection was successful in identifying a sub‐cohort that will, on average, have a higher rate of decline than the rest. The proposed identification of decliners has the potential of reducing trial costs by minimizing the enrolment of cognitively stable individuals into trials. Reference: (1) The ADNI1 and ADNI2 datasets were obtained from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database (adni.loni.usc.edu).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.209 | 0.239 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".