Validation of ICD‐10 diagnosis codes for identification of veterans with intracerebral hemorrhage and subarachnoid hemorrhage using clinical notes in the United States Veterans Affairs Healthcare System
Bibliographic record
Abstract
Abstract Background Cerebral amyloid angiopathy (CAA) is a significant contributor to hemorrhagic stroke, notably lobar intracerebral hemorrhage (ICH) and convexity subarachnoid hemorrhage (SAH). This study describes the natural occurrence of ICH and SAH events among veterans, including those with AD, within the United States Veterans Affairs Healthcare System (VAHS). Method The VAHS database was evaluated to identify ICD‐10 codes for ICH (I61.x) and SAH (I60.x) from 2015‐2023. A subsample of veterans with AD was identified based on 1 qualifier (AD diagnostic code or clinical note); a sensitivity analysis included veterans with ≥2 AD qualifiers, ≥30 days apart. Two‐thousand veterans were randomly selected from the ICH/SAH sample for validation of diagnostic coding using clinical notes. The positive predictive value (PPV) of ICH/SAH diagnostic codes was determined using notes from 100 randomly selected cases. Result A total of 23,539 and 7,822 veterans were identified using ICD‐10 codes for ICH and SAH, respectively, out of 4‐5 million veterans receiving care annually. The ICH/SAH sample was 93/95% male, 65/69% white, with a mean age of 70 years. Approximately 14% and 4/5% of veterans with ICH/SAH had AD based on 1 and 2 AD qualifiers, respectively. From 2016‐2023, the yearly prevalence of ICH and SAH was approximately 8‐10 and 3/10,000 patients, respectively ( Figure ). Approximately 61/68% of veterans in the ICH/SAH sample were identified from outpatient visits only and 39/32% were identified from VAHS inpatient stays with or without outpatient medical records. Among 2,000 randomly selected cases for coding validation, 98% had notes available within ±14 days of the ICD coding; among these, 55/60% carried an ICH/SAH keyword. Brain MRI records were found in one‐third of ICH/SAH cases. Review of 2,000 clinical notes corresponding to 100 randomly selected cases found documentation of ICH in 80% (PPV) of veterans with an ICH diagnostic code and 100% with an SAH diagnostic code. Conclusion Yearly prevalence of ICH/SAH (2016‐2023) was 0.08‐0.1%/0.03% in US veterans with 14% of the total cases carrying ≥1 AD identifier. Among cases with clinical notes for ICH and SAH, PPVs were 80% and 100%, respectively.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".