Using real‐world general practitioner data to study the diagnosis and management of dementia: rationale and design
Bibliographic record
Abstract
Abstract Background General practitioners (GPs) play a critical role in the early recognition of cognitive deficits and management of dementia. Timely diagnosis is important in light of potential disease‐modifying therapies and the potential to improve patient outcomes. We aimed to establish a real‐world data cohort utilizing data from GPs on individuals with dementia as a starting point, with the goal of gaining valuable insights into trajectories and management of dementia and patient outcomes. Here we describe the rationale and design. Method We selected individuals with dementia using Dutch GP data from the PHARMO Data Network, which includes diagnoses, symptoms, examinations, prescriptions, and communication between GPs and specialists. Diagnosis of dementia was defined as a diagnosis code or prescription of anti‐dementia drugs between January 1, 2011 and December 31, 2020. Persons were included if they had one year of history prior to dementia diagnosis. We described the cohort in terms of demographics and screening tests for cognitive impairment. Result A total of 52,911 individuals with dementia were selected from a source population of 4.7 million persons. The mean age was 81 years (standard deviation [SD] = 8.65) and 31,343 (59%) were female (Table 1). On average, patients have 8.6 years (SD = 4.09) of data available prior to dementia diagnosis and can be followed‐up for 2.8 years (SD = 2.19) after diagnosis. The reason for end of follow‐up was death for 16,978 persons (32%), end of data availability for 21,314 persons (40%), and 14,699 (28%) reached December 31, 2020 and are still registered (i.e., active). We found the Mini‐Mental State Examination (MMSE) in GP records of 31,759 persons (60%), and the Montreal Cognitive Assessment (MoCA) and Rowland Universal Dementia Assessment Scale (RUDAS) for only 1,777 (3%), and 29 persons (<0.5%), respectively. Conclusion We created a cohort of 52,991 individuals with dementia, providing a starting point for further research on trajectories and management of AD in primary care and patient outcomes. Next steps include matching the cohort with dementia‐free controls, enriching the cohort by established linkages to other data sources (e.g. hospital data), and examining healthcare resource utilization, indicators of cognitive decline, treatment, and young‐onset dementia.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.103 | 0.152 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.006 | 0.008 |
| Science and technology studies | 0.003 | 0.004 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.004 | 0.005 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".