Uncovering Alzheimer's Disease Related Dementias Patterns using Machine Learning in a LMIC Setting
Bibliographic record
Abstract
Abstract Background Neurodegenerative disease prevalence is projected to triple in the next 30 years, with sub‐Saharan Africa facing an acute brain health crisis due to population aging and a high vascular risk burden. East Africa, where 80% of the population is under 35, is expected to experience rapid brain aging and the world's highest neurological disease burden in terms of disability‐adjusted life‐years. Despite this, research on brain aging specific to this population is limited. Machine learning (ML) models trained on magnetic resonance imaging (MRI) scans offer promising diagnostic and prognostic insights for early detection of Alzheimer's disease and related dementias (ADRD). This study aims to develop an ensemble ML model tailored to improve early diagnosis. Methods We will curate a neuroimaging database (Kenya‐NeuroBank) consisting of T1‐weighted MRI scans, demographic, and clinical data from ADRD diagnostic services at the Aga Khan University Hospital, Nairobi. This dataset, encompassing neurotypical and ADRD‐affected brains, will be used to train an ensemble ML model for ADRD prediction in the Kenyan population. Additionally, we will identify key brain structures most associated with ADRD using ML‐based feature selection techniques. Model validation will be conducted using complementary clinical datasets from Kenya and available datasets from the African American cohort (such as Wake Forest Alzheimer's Disease Research Center data and HABS‐HD study data), enabling comparisons across populations. Results Currently, as we curate the Kenya‐NeuroBank database, 136 participant data is available. Of these 67.6% ( n = 92) are females and the mean age is 51.7 years (SD=11.4). We anticipate creating the database to around 1000 participants. We expect our ML model to accurately classify ADRD cases, identify population‐specific structural biomarkers, and improve early detection of neurodegenerative risks in East Africa. Comparative validation with Western datasets will highlight unique environmental and lifestyle influences on ADRD pathology. Conclusion This study addresses a critical gap in neuroimaging research in sub‐Saharan Africa, providing a tailored ML approach for ADRD prediction. By improving early diagnosis and identifying population‐specific biomarkers, our findings will contribute to personalized healthcare strategies, earlier interventions, and better patient outcomes. This work also has broader implications for brain health research across Africa.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".