A Stratified Machine Learning Evaluation of Risk Factors of Dementia Conversion
Bibliographic record
Abstract
BACKGROUND: Understanding the complex interplay of risk factors for dementia is essential for developing effective prevention strategies. Older adults present a high frequency of multimorbidity, though risk factors of dementia are usually evaluated individually in this population. In this study, we aim to simultaneously identify modifiable and non-modifiable risk factors in predicting dementia conversion. METHOD: Real-world longitudinal data from the National Alzheimer's Coordinating Center (NACC), spanning 2005 to 2023 across 46 Alzheimer's Disease Research Centers (ADRCs) was analysed. Eleven modifiable risk factors were stratified: hearing loss, hypertension, body mass index (BMI), depression, visual loss, education, hyperlipidemia, traumatic brain injury (TBI), alcohol abuse, smoking, and diabetes. Age and gender were analyzed as non-modifiable factors. A machine learning approach was employed for simultaneous evaluation of risk factors (Figure 1). RESULT: We included 11,107 cognitively unimpaired individuals at baseline, whose 1,052 converted to dementia (Tab. 1). SHAP (SHapley Additive exPlanations) value analysis assessed the impact of each factor, with Figure 2a displaying the distribution of impacts and Figure 2b illustrating their absolute contributions. The overall performance of the model included an average accuracy of predicting dementia conversion of 66.03% (95% CI: [65.82, 66.23]), sensitivity of 70.26% (95% CI: [70.03, 70.48]), specificity of 65.59% (95% CI: [65.34, 65.83]), and an ROC-AUC of 0.745 (95% CI: [0.744, 0.745]). Age was the most impactful, with an individual AUC score of 0.724 (Figure 2c), and a strong influence on model performance (Figure 2d). Hearing loss, hypertension, BMI, depression, visual loss and education emerged as impactful modifiable risk-factors in our model, albeit not as much as age. CONCLUSION: This study highlights the significance of a simultaneous evaluation of risk factors of dementia, considering multimorbidities in dementia prevention. While age is a primary predictor, our approach identifies critical points within modifiable factors like hearing loss and hypertension. The machine learning framework enhances predictive accuracy, offering comprehensive insights for prevention strategies. Future studies should focus on validating these findings across diverse populations and longitudinally.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.018 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".