Coding of Childhood Psychiatric and Neurodevelopmental Disorders in Electronic Health Records of a Large Integrated Health Care System: Validation Study
Notice bibliographique
Résumé
Background: Mental, emotional, and behavioral disorders are chronic pediatric conditions, and their prevalence has been on the rise over recent decades. Affected children have long-term health sequelae and a decline in health-related quality of life. Due to the lack of a validated database for pharmacoepidemiological research on selected mental, emotional, and behavioral disorders, there is uncertainty in their reported prevalence in the literature. objectives: We aimed to evaluate the accuracy of coding related to pediatric mental, emotional, and behavioral disorders in a large integrated health care system's electronic health records (EHRs) and compare the coding quality before and after the implementation of the International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) coding as well as before and after the COVID-19 pandemic. Methods: Medical records of 1200 member children aged 2-17 years with at least 1 clinical visit before the COVID-19 pandemic (January 1, 2012, to December 31, 2014, the ICD-9-CM coding period; and January 1, 2017, to December 31, 2019, the ICD-10-CM coding period) and after the COVID-19 pandemic (January 1, 2021, to December 31, 2022) were selected with stratified random sampling from EHRs for chart review. Two trained research associates reviewed the EHRs for all potential cases of autism spectrum disorder (ASD), attention-deficit hyperactivity disorder (ADHD), major depression disorder (MDD), anxiety disorder (AD), and disruptive behavior disorders (DBD) in children during the study period. Children were considered cases only if there was a mention of any one of the conditions (yes for diagnosis) in the electronic chart during the corresponding time period. The validity of diagnosis codes was evaluated by directly comparing them with the gold standard of chart abstraction using sensitivity, specificity, positive predictive value, negative predictive value, the summary statistics of the F-score, and Youden J statistic. κ statistic for interrater reliability among the 2 abstractors was calculated. Results: The overall agreement between the identification of mental, behavioral, and emotional conditions using diagnosis codes compared to medical record abstraction was strong and similar across the ICD-9-CM and ICD-10-CM coding periods as well as during the prepandemic and pandemic time periods. The performance of AD coding, while strong, was relatively lower compared to the other conditions. The weighted sensitivity, specificity, positive predictive value, and negative predictive value for each of the 5 conditions were as follows: 100%, 100%, 99.2%, and 100%, respectively, for ASD; 100%, 99.9%, 99.2%, and 100%, respectively, for ADHD; 100%, 100%, 100%, and 100%, respectively for DBD; 87.7%, 100%, 100%, and 99.2%, respectively, for AD; and 100%, 100%, 99.2%, and 100%, respectively, for MDD. The F-score and Youden J statistic ranged between 87.7% and 100%. The overall agreement between abstractors was almost perfect (κ=95%). Conclusions: Diagnostic codes are quite reliable for identifying selected childhood mental, behavioral, and emotional conditions. The findings remained similar during the pandemic and after the implementation of the ICD-10-CM coding in the EHR system.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,020 | 0,065 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».