A218 VALIDATION OF A CIRRHOSIS CASE DEFINITION IN CANADIAN ADMINISTRATIVE DATA
Bibliographic record
Abstract
Administrative and population-based databases can be utilized for health services research related to cirrhosis. The performance of coding algorithms for the identification of patients with cirrhosis in Canadian administrative data has not been previously defined. To validate the use of International Classification of Disease (ICD) algorithms to identify patients with cirrhosis using administrative data from Ontario, Canada. We performed primary chart abstraction of 458 consecutive patients seen in the tertiary care Liver Clinic at Hotel Dieu Hospital in Kingston, Ontario from May – August 2013. In order to define the presence or absence of cirrhosis and related decompensations, details regarding the etiology and severity of liver disease were abstracted. The gold-standard definition of cirrhosis was based on the presence of cirrhosis after chart review by two hepatologists. This data was then linked to the administrative databases of the Institute for Clinical Evaluative Sciences. We used ICD-9, ICD-10 and Ontario Health Insurance Plan billing codes for cirrhosis, cirrhotic decompensations, and chronic liver diseases to develop multiple coding algorithms for the identification of cirrhosis. The sensitivity, specificity, positive predictive value (PPV) and negative predictive value (NPV) were calculated for each based on the gold standard cirrhosis definition. A total of 10 different algorithms were evaluated. Overall, the use of one inpatient or one outpatient code for cirrhosis resulted in the highest sensitivity (79%; CI 74% - 84%) with a specificity of 79% (CI: 73% - 85%), PPV 81% (CI: 75%-86%), and NPV 78% (CI: 71%-83%). Using 2 outpatient or 1 inpatient codes for cirrhosis plus a decompensation code plus a chronic liver disease code resulted in the highest specificity (99%; CI: 96% - 100%) and PPV 95% (CI: 86% - 99%), however this was associated with a large drop-off in sensitivity (24%; CI: 19% - 31%) and NPV 54% (CI: 49% - 59%). The use of ICD coding algorithms can effectively identify patients with cirrhosis using administrative data from Canada and can be used in future health services research studies. Southeastern Ontario Academic Medical Association New Clinician Scientist Award
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.035 | 0.093 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".