Validity of an algorithm to identify cardiovascular deaths from administrative health records: a multi-database population-based cohort study
Bibliographic record
Abstract
BACKGROUND: Cardiovascular death is a common outcome in population-based studies about new healthcare interventions or treatments, such as new prescription medications. Vital statistics registration systems are often the preferred source of information about cause-specific mortality because they capture verified information about the deceased, but they may not always be accessible for linkage with other sources of population-based data. We assessed the validity of an algorithm applied to administrative health records for identifying cardiovascular deaths in population-based data. METHODS: Administrative health records were from an existing multi-database cohort study about sodium-glucose cotransporter-2 (SGLT2) inhibitors, a new class of antidiabetic medications. Data were from 2013 to 2018 for five Canadian provinces (Alberta, British Columbia, Manitoba, Ontario, Quebec) and the United Kingdom (UK) Clinical Practice Research Datalink (CPRD). The cardiovascular mortality algorithm was based on in-hospital cardiovascular deaths identified from diagnosis codes and select out-of-hospital deaths. Sensitivity, specificity, and positive and negative predictive values (PPV, NPV) were calculated for the cardiovascular mortality algorithm using vital statistics registrations as the reference standard. Overall and stratified estimates and 95% confidence intervals (CIs) were computed; the latter were produced by site, location of death, sex, and age. RESULTS: The cohort included 20,607 individuals (58.3% male; 77.2% ≥70 years). When compared to vital statistics registrations, the cardiovascular mortality algorithm had overall sensitivity of 64.8% (95% CI 63.6, 66.0); site-specific estimates ranged from 54.8 to 87.3%. Overall specificity was 74.9% (95% CI 74.1, 75.6) and overall PPV was 54.5% (95% CI 53.7, 55.3), while site-specific PPV ranged from 33.9 to 72.8%. The cardiovascular mortality algorithm had sensitivity of 57.1% (95% CI 55.4, 58.8) for in-hospital deaths and 72.3% (95% CI 70.8, 73.9) for out-of-hospital deaths; specificity was 88.8% (95% CI 88.1, 89.5) for in-hospital deaths and 58.5% (95% CI 57.3, 59.7) for out-of-hospital deaths. CONCLUSIONS: A cardiovascular mortality algorithm applied to administrative health records had moderate validity when compared to vital statistics data. Substantial variation existed across study sites representing different geographic locations and two healthcare systems. These variations may reflect different diagnostic coding practices and healthcare utilization patterns.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.201 | 0.416 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.006 | 0.003 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".