Applying machine learning methods to describe complex medication use in a population of community-dwelling older adults living with dementia: lessons for pharmacoepidemiology studies using health administrative data
Bibliographic record
Abstract
Persons living with dementia often have complex medication use due to symptoms related to dementia and comorbidities that accumulate as individuals age. Understanding the consequences of using multiple medications in combination, that is, the potential risk of interactions between medications being used concurrently, is important to optimizing medication use, and reducing related adverse events. Unsupervised machine learning methods offer an opportunity to incorporate more data about exposure to multiple medications into pharmacoepidemiology methods to better understand prescribing patterns and medication related adverse events. This thesis examined the application of two unsupervised machine learning methods – network analysis and hierarchical clustering – to describe polypharmacy and fall-related hospitalizations over time in a population-based cohort of community-dwelling older adults with incident dementia in three related studies in linked health administrative data. In Paper One, network analysis described the common medication subclasses concurrently prescribed within persons living with dementia at case ascertainment and five years following. In Paper Two, hierarchical clustering found groups based on medication use were associated with comorbidities. In Paper Three, there was an association between the CNS-active medication prescribing cluster (individuals with similar, higher-than-average CNS-active medication use) and fall-related hospitalizations, a potentially medication-related adverse event, even when controlling for sex, age, potentially inappropriate prescribing, and level of polypharmacy. Collectively, the results from these studies demonstrate the benefits (and limitations) of unsupervised machine learning methods – including hypothesis generation, data reduction, and summarizing complex information visually – in pharmacoepidemiology studies. From a population health perspective, in persons living with dementia, this thesis provided an overview of prescribing patterns, demonstrated the importance of cardiovascular and CNS-active medications (particularly the latter’s association with falls), and highlighted the need for proper management and treatment of comorbidities alongside the management of dementia.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.041 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".