Comparison of Coding of Heart Failure and Comorbidities in Administrative and Clinical Data for Use in Outcomes Research
Bibliographic record
Abstract
BACKGROUND: Despite the potential usefulness of administrative databases for evaluating outcomes, coding of heart failure and associated comorbidities have not been definitively compared with clinical data. OBJECTIVE: To compare the predictive value of heart failure diagnoses and secondary conditions identified in a large administrative database with chart-based records. METHODS: The authors studied 1808 patient records sampled from 14 acute care hospitals and compared clinically recorded data with administrative records from the Canadian Institute for Health Information. The impact of comorbidity coding in the administrative data set according to the Charlson classification was examined in models of 30-day mortality. RESULTS: The positive predictive value (PPV) of a primary diagnosis ICD-9 428 was 94.3% using the Framingham criteria and 88.6% using criteria previously validated with pulmonary capillary wedge pressure. There was reduced prevalence of secondary comorbid conditions in administrative data in comparison with clinical chart data. The specificities and PPV/negative predictive values of administratively identified index comorbidities were high. The sensitivities of index comorbidities were low, but were enhanced by examination of hospitalizations within 1 year prior to the index heart failure admission. Using information from prior hospitalizations modestly enhanced 30-day mortality model performance; however, the odds ratio point estimates of the index and enhanced administrative data sets were consistent with the clinical model. CONCLUSION: The ICD-9 428 primary diagnosis is highly predictive of heart failure using clinical criteria. Examination of hospitalization data up to 1 year prior to the index admission improves comorbidity detection and may provide enhancements to future studies of heart failure mortality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.121 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".