Positive predictive value and sensitivity of ICD‐9‐CM codes for identifying pediatric leukemia
Bibliographic record
Abstract
BACKGROUND: To facilitate community-based epidemiologic studies of pediatric leukemia, we validated use of ICD-9-CM diagnosis codes to identify pediatric leukemia cases in electronic medical records of six U.S. integrated health plans from 1996-2015 and evaluated the additional contributions of procedure codes for diagnosis/treatment. PROCEDURES: Subjects (N = 408) were children and adolescents born in the health systems and enrolled for at least 120 days after the date of the first leukemia ICD-9-CM code or tumor registry diagnosis. The gold standard was the health system tumor registry and/or medical record review. We calculated positive predictive value (PPV) and sensitivity by number of ICD-9-CM codes received in the 120-day period following and including the first code. We evaluated whether adding chemotherapy and/or bone marrow biopsy/aspiration procedure codes improved PPV and/or sensitivity. RESULTS: Requiring receipt of one or more codes resulted in 99% sensitivity (95% confidence interval [CI]: 98-100%) but poor PPV (70%; 95% CI: 66-75%). Receipt of two or more codes improved PPV to 90% (95% CI: 86-93%) with 96% sensitivity (95% CI: 93-98%). Requiring at least four codes maximized PPV (95%; 95% CI: 92-98%) without sacrificing sensitivity (93%; 95% CI: 89-95%). Across health plans, PPV for four codes ranged from 84-100% and sensitivity ranged from 83-95%. Including at least one code for a bone marrow procedure or chemotherapy treatment had minimal impact on PPV or sensitivity. CONCLUSIONS: The use of diagnosis codes from the electronic health record has high PPV and sensitivity for identifying leukemia in children and adolescents if more than one code is required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.064 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".