Predicting progression in subjective cognitive decline (SCD) using a machine learning (ML) approach: The role of the complaint’s severity
Bibliographic record
Abstract
Abstract Background Presence of significant subjective complaints about cognition (SCD) is considered the first behavioral manifestation of Alzheimer disease (AD). However, SCD has not yet overcome the challenge of becoming a reliable preclinical AD marker. Severity indices were proposed to improve the accuracy of complaints when predicting the risk of AD (Jessen et al., 2010). Our aim was to compare the predictive accuracy of ML algorithms using a more (95%ile) or less (5%ile) restrictive cut‐off point in severity of complaints for classification in Low (LSC) and High (HSC) subjective complaints groups. Method One hundred and ninety‐nine participants from the Compostela Aging Study (ComPAS) were identified at baseline as SCDs and completed three follow‐up measurements (54‐72 months). Considering their total scores in the Subjective Memory Questionnaire, participants were classified as LSC and HSC following using two distinct cut‐off points: complaint scores above or below than 5%ile vs 95%ile. Participants were labeled as ‘worsening’ or ‘stable’ based on their progression or not to MCI or AD. ML classifier algorithms (Random Forest, Support Vector Machine, Extra Tree) were applied to forty‐one measures (socio‐demographic, time, health, cognitive, behavioral, cognitive reserve) collected at baseline. Results The best performing model was the Random Forest (95%ile: PPV =.87; Sensibility =.92; Specificity =.38; 5%ile: PPV =.66; Sensibility=.61; Specificity =.51). The confusion matrix (Figure 1) showed that: a) both, the more (95%ile) and the less (5%ile) restrictive criteria mostly classified the HSC‐stable participants as LSC‐stable; and b) criterion 5%ile, but not 95%ile, was able to differentially identify progressors based on their complaints. Episodic memory, executive functions, depression, cognitive reserve and progression timing measures assume the highest levels of importance in the ML algorithm that differentially predicts the progression of SCD to MCI and AD according to the level of complaints. Conclusions For both the criteria, algorithms failed to successfully identify stable SCD participants. The less restrictive criterion resulted in a better classification than the more lenient one to differentially identify LSC and HSC participants who get worse. Cognitive, affective, cognitive reserve and progression timing were the more prominent variables in conforming the predictive algorithm.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".