Postpartum Migraine Headache Coding in Electronic Health Records of a Large Integrated Health Care System: Validation Study
Bibliographic record
Abstract
BACKGROUND: Migraine is a common neurological disorder characterized by repeated headaches of varying intensity. The prevalence and severity of migraine headaches disproportionally affects women, particularly during the postpartum period. Moreover, migraines during pregnancy have been associated with adverse maternal outcomes, including preeclampsia and postpartum stroke. However, due to the lack of a validated instrument for uniform case ascertainment on postpartum migraine headache, there is uncertainty in the reported prevalence in the literature. OBJECTIVE: The aim of this study was to evaluate the completeness and accuracy of reporting postpartum migraine headache coding in a large integrated health care system's electronic health records (EHRs) and to compare the coding quality before and after the implementation of the International Classification of Diseases, 10th revision, Clinical Modification (ICD-10-CM) codes and pharmacy records in EHRs. METHODS: Medical records of 200 deliveries in all 15 Kaiser Permanente Southern California hospitals during 2 time periods, that is, January 1, 2012 through December 31, 2014 (International Classification of Diseases, 9th revision, Clinical Modification [ICD-9-CM] coding period) and January 1, 2017 through December 31, 2019 (ICD-10-CM coding period), were randomly selected from EHRs for chart review. Two trained research associates reviewed the EHRs for all 200 women for postpartum migraine headache cases documented within 1 year after delivery. Women were considered to have postpartum migraine headache if either a mention of migraine headache (yes for diagnosis) or a prescription for treatment of migraine headache (yes for pharmacy records) was noted in the electronic chart. Results from the chart abstraction served as the gold standard and were compared with corresponding diagnosis and pharmacy prescription utilization records for both ICD-9-CM and ICD-10-CM coding periods through comparisons of sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), as well as the summary statistics of F-score and Youden J statistic (J). The kappa statistic (κ) for interrater reliability was calculated. RESULTS: The overall agreement between the identification of migraine headache using diagnosis codes and pharmacy records compared to the medical record review was strong. Diagnosis coding (F-score=87.8%; J=82.5%) did better than pharmacy records (F-score=72.7%; J=57.5%) when identifying cases, but combining both of these sources of data produced much greater accuracy in the identification of postpartum migraine cases (F-score=96.9%; J=99.7%) with sensitivity, specificity, PPV, and NPV of 100%, 99.7%, 93.9%, and 100%, respectively. Results were similar across the ICD-9-CM (F-score=98.7%, J=99.9%) and ICD-10-CM coding periods (F-score=94.9%; J=99.6%). The interrater reliability between the 2 research associates for postpartum migraine headache was 100%. CONCLUSIONS: Neither diagnostic codes nor pharmacy records alone are sufficient for identifying postpartum migraine cases reliably, but when used together, they are quite reliable. The completeness of the data remained similar after the implementation of the ICD-10-CM coding in the EHR system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.055 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".