Limitations of pulmonary embolism ICD-10 codes in emergency department administrative data: let the buyer beware
Bibliographic record
Abstract
BACKGROUND: Administrative data is a useful tool for research and quality improvement; however, validity of research findings based on these data depends on their reliability. Diagnoses assigned by physicians are subsequently converted by nosologists to ICD-10 codes (International Statistical Classification of Diseases and Related Health Problems, 10th Revision). Several groups have reported ICD-9 coding errors in inpatient data that have implications for research, quality improvement, and policymaking, but few have assessed ICD-10 code validity in ambulatory care databases. Our objective was to evaluate pulmonary embolism (PE) ICD-10 code accuracy in our large, integrated hospital system, and the validity of using these codes for operational and health services research using ED ambulatory care databases. METHODS: Ambulatory care data for patients (age ≥ 18 years) with a PE ICD-10 code (I26.0 and I26.9) were obtained from the records of four urban EDs between July 2013 to January 2015. PE diagnoses were confirmed by reviewing medical records and imaging reports. In cases where chart diagnosis and ICD-10 code were discrepant, chart review was considered correct. Physicians' written discharge diagnoses were also searched using 'pulmonary embolism' and 'PE', and patients who were diagnosed with PE but not coded as PE were identified. Coding discrepancies were quantified and described. RESULTS: One thousand, four hundred and fifty-three ED patients had a PE ICD-10 code. Of these, 257 (17.7%) were false positive, with an incorrectly assigned PE code. Among the 257 false positives, 193 cases had ambiguous ED diagnoses such as 'rule out PE' or 'query PE', while 64 cases should have had non-PE codes. An additional 117 patients (8.90%) with a PE discharge diagnosis were incorrectly assigned a non-PE ICD-10 code (false negative group). The sensitivity of PE ICD-10 codes in this dataset was 91.1% (95%CI, 89.4-92.6) with a specificity of 99.9% (95%CI, 99.9-99.9). The positive and negative predictive values were 82.3% (95%CI, 80.3-84.2) and 99.9% (95%CI, 99.9-99.9), respectively. CONCLUSIONS: Ambulatory care data, like inpatient data, are subject to coding errors. This confirms the importance of ICD-10 code validation prior to use. The largest proportion of coding errors arises from ambiguous physician documentation; therefore, physicians and data custodians must ensure that quality improvement processes are in place to promote ICD-10 coding accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.117 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".