Identifying patterns of co-occurring chronic conditions preceding dementia: An unsupervised machine learning approach using health administrative data
Bibliographic record
Abstract
ObjectivesIndividual risk factors for dementia are well known, but the influence of co-occurring chronic conditions has not been considered. We identified clusters of chronic conditions using an unsupervised machine learning approach and examined associations with incident dementia. ApproachUsing linked population-based administrative databases, we followed all community-dwelling adults aged 40-54 years in Ontario, Canada from April 2002 until March 2019 for incident dementia. We estimated the prevalence of 29 chronic conditions using validated algorithms and/or diagnosis codes. We reduced dataset dimensionality using multiple correspondence analysis and a fuzzy c-means clustering algorithm identified the optimal number of clusters (between 3-6 tested). Associations between clusters and incident dementia were examined using a cause-specific hazard model adjusted for sociodemographic characteristics and accounting for the competing risk of death. ResultsWe identified 82,359 eligible individuals (random 3% sample of total eligible individuals; mean age 46.5 years; 50.4% female). Regression analyses were based on 5 comorbidity clusters (fuzzy silhouette index:0.69). Compared to the low comorbidity cluster, persons in the cerebrovascular disease/metabolic (HRadj=3.06, 95%CI[2.42,3.86]) and neuro-related/mental health clusters (HRadj=2.51, 95%CI[2.05,3.07]) had the highest rates of incident dementia, followed by the cardiovascular risk factor cluster (HRadj=1.66,95%CI[1.32,2.09]). Persons in the cancer cluster did not have an increased incidence of dementia (HRadj=0.96,95%CI[0.77,1.20]). ConclusionsWe found significant associations between machine learning-derived clusters of chronic conditions and dementia. ImplicationsUnsupervised machine learning approaches to identify clusters of chronic conditions may be a useful tool for considering the impact of multimorbidity on dementia risk.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.004 | 0.000 |
| Scholarly communication | 0.001 | 0.006 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".