Coding of Childhood Psychiatric and Neurodevelopmental Disorders in Electronic Health Records of a Large Integrated Health Care System: Validation Study
Bibliographic record
Abstract
Background: Mental, emotional, and behavioral disorders are chronic pediatric conditions, and their prevalence has been on the rise over recent decades. Affected children have long-term health sequelae and a decline in health-related quality of life. Due to the lack of a validated database for pharmacoepidemiological research on selected mental, emotional, and behavioral disorders, there is uncertainty in their reported prevalence in the literature. objectives: We aimed to evaluate the accuracy of coding related to pediatric mental, emotional, and behavioral disorders in a large integrated health care system's electronic health records (EHRs) and compare the coding quality before and after the implementation of the International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) coding as well as before and after the COVID-19 pandemic. Methods: Medical records of 1200 member children aged 2-17 years with at least 1 clinical visit before the COVID-19 pandemic (January 1, 2012, to December 31, 2014, the ICD-9-CM coding period; and January 1, 2017, to December 31, 2019, the ICD-10-CM coding period) and after the COVID-19 pandemic (January 1, 2021, to December 31, 2022) were selected with stratified random sampling from EHRs for chart review. Two trained research associates reviewed the EHRs for all potential cases of autism spectrum disorder (ASD), attention-deficit hyperactivity disorder (ADHD), major depression disorder (MDD), anxiety disorder (AD), and disruptive behavior disorders (DBD) in children during the study period. Children were considered cases only if there was a mention of any one of the conditions (yes for diagnosis) in the electronic chart during the corresponding time period. The validity of diagnosis codes was evaluated by directly comparing them with the gold standard of chart abstraction using sensitivity, specificity, positive predictive value, negative predictive value, the summary statistics of the F-score, and Youden J statistic. κ statistic for interrater reliability among the 2 abstractors was calculated. Results: The overall agreement between the identification of mental, behavioral, and emotional conditions using diagnosis codes compared to medical record abstraction was strong and similar across the ICD-9-CM and ICD-10-CM coding periods as well as during the prepandemic and pandemic time periods. The performance of AD coding, while strong, was relatively lower compared to the other conditions. The weighted sensitivity, specificity, positive predictive value, and negative predictive value for each of the 5 conditions were as follows: 100%, 100%, 99.2%, and 100%, respectively, for ASD; 100%, 99.9%, 99.2%, and 100%, respectively, for ADHD; 100%, 100%, 100%, and 100%, respectively for DBD; 87.7%, 100%, 100%, and 99.2%, respectively, for AD; and 100%, 100%, 99.2%, and 100%, respectively, for MDD. The F-score and Youden J statistic ranged between 87.7% and 100%. The overall agreement between abstractors was almost perfect (κ=95%). Conclusions: Diagnostic codes are quite reliable for identifying selected childhood mental, behavioral, and emotional conditions. The findings remained similar during the pandemic and after the implementation of the ICD-10-CM coding in the EHR system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.065 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".