MétaCan
Menu
Back to cohort
Record W3127275001 · doi:10.1016/j.jid.2021.01.009

Accuracy of Algorithms to Identify People with Atopic Dermatitis in Ontario Routinely Collected Health Databases

2021· article· en· W3127275001 on OpenAlexafffundabout
Mohamed Abdalla, Branson Chen, Robin Santiago, Jacqueline Young, Lihi Eder, An‐Wen Chan, Elena Pope, Karen Tu, Liisa Jaakkimainen, Aaron M. Drucker

Bibliographic record

VenueJournal of Investigative Dermatology · 2021
Typearticle
Languageen
FieldMedicine
TopicDermatology and Skin Diseases
Canadian institutionsHealth Sciences CentreSunnybrook Health Science CentreNorth York General HospitalToronto Western HospitalHospital for Sick ChildrenInstitute for Clinical Evaluative SciencesVector InstituteWomen's College HospitalUniversity Health NetworkUniversity of Toronto
FundersCanadian Institutes of Health ResearchOntario Ministry of Health and Long-Term CarePhysicians' Services Incorporated FoundationInstitute for Clinical Evaluative SciencesCanadian Dermatology Foundation
KeywordsAtopic dermatitisDatabaseMedicineAlgorithmComputer scienceDermatology

Abstract

fetched live from OpenAlex

Atopic dermatitis (AD) is associated with high patient and population burden (Bridgman et al., 2018Bridgman A.C. Block J.K. Drucker A.M. The multidimensional burden of atopic dermatitis: an update.Ann Allergy Asthma Immunol. 2018; 120: 603-606Abstract Full Text Full Text PDF PubMed Scopus (23) Google Scholar) and increased healthcare resource utilization (Drucker et al., 2018Drucker A.M. Qureshi A.A. Amand C. Villeneuve S. Gadkari A. Chao J. et al.Health care resource utilization and costs among adults with atopic dermatitis in the United States: a claims-based analysis.J Allergy Clin Immunol Pract. 2018; 6: 1342-1348Abstract Full Text Full Text PDF PubMed Scopus (26) Google Scholar). These factors make understanding the epidemiology of AD and associated health service utilization a priority. Routinely collected data, including electronic medical records and health administrative data, are valuable resources, but the accuracy of case definitions to identify people with AD in these databases must be assessed to ensure validity (Dizon et al., 2018Dizon M.P. Yu A.M. Singh R.K. Wan J. Chren M.M. Flohr C. et al.Systematic review of atopic dermatitis disease definition in studies using routinely collected health data.Br J Dermatol. 2018; 178: 1280-1287Crossref PubMed Scopus (22) Google Scholar). Our objective was to test case definitions of AD in routinely collected health data for Ontario, Canada. We followed reporting recommendations for studies assessing accuracy of diagnoses in health administrative data (Benchimol et al., 2011Benchimol E.I. Manuel D.G. To T. Griffiths A.M. Rabeneck L. Guttmann A. Development and use of reporting guidelines for assessing the quality of validation studies of health administrative data.J Clin Epidemiol. 2011; 64: 821-829Abstract Full Text Full Text PDF PubMed Scopus (211) Google Scholar). We derived our testing population from the electronic medical record primary care database (EMRPC, also known as EMRALD), which is linked with health administrative data for Ontario, Canada (Supplementary Materials and Methods). EMRPC contains clinical notes, problem lists, prescribed medications, and specialist consultation notes for 43 family practice clinics with 367,166 rostered patients, with physician and patient characteristics representative of the general population (Tu et al., 2015Tu K. Widdifield J. Young J. Oud W. Ivers N.M. Butt D.A. et al.Are family physicians comprehensively using electronic medical records such that the data can be used for secondary purposes? A Canadian perspective.BMC Med Inform Decis Mak. 2015; 15: 67Crossref PubMed Scopus (35) Google Scholar). We performed an electronic search for “eczema” and “dermatitis” in an age-stratified random sample of 7,268 EMRPC charts (2,036 children <18 years old, 5,232 adults ≥18 years old) from 1987 to 2016, which served as the testing population (Figure 1 and Supplementary Table S1). One of three trained abstractors, blinded to administrative data, reviewed each chart containing those keywords to determine whether patients had AD, defined by use of the word “atopic” preceding “eczema” or “dermatitis” or by features satisfying modified United Kingdom Working Party diagnostic criteria (Williams et al., 1994Williams H.C. Burney P.G. Pembroke A.C. Hay R.J. The UK. Working Party’s Diagnostic Criteria for Atopic Dermatitis. III. Independent hospital validation.Br J Dermatol. 1994; 131: 406-416Crossref PubMed Scopus (679) Google Scholar) (Supplementary Materials and Methods). When abstractors were unsure how to classify a case, it was adjudicated by a dermatologist (AMD). Duplicate abstraction was performed on 118 charts with good agreement (112 of 118, 95%) and interrater reliability (Cohen’s k = 0.67); discrepant results were adjudicated by AMD. We tested the following different means of identifying people with AD in the testing population:1.rules-based simple text mining of clinical notes and structured electronic medical record variables for keywords, diagnostic billing codes, and treatments associated with AD (human-defined features) within EMRPC patient charts (e.g., presence of AD in clinical notes or problem list or International Classification of Diseases 9 code 691 in billed codes);2.multiple machine learning (ML) algorithms on the same human-defined features as the first points within EMRPC patient charts (see Supplementary Materials and Methods for details);3.rules-based combinations of International Classification of Diseases codes within health administrative data;4.multiple ML algorithms on health administrative data. For the ML algorithms using human-defined features within EMRPC charts, we experimented using different classifiers, including logistic regression, random forests, and shallow neural networks. For free-text ML classification (i.e., using complete notes instead of human-defined features), we trained various ML classifiers (e.g., support vector machine, logistic regression, shallow neural network) on numeric representations of free-text fields using the tf-idf algorithm (Mohammadhassanzadeh et al., 2020Mohammadhassanzadeh H. Sketris I. Traynor R. Alexander S. Winquist B. Stewart S.A. Using natural language processing to examine the uptake, content, and readability of media coverage of a pan-Canadian drug safety research project: cross-sectional observational study.JMIR Form Res. 2020; 4: e13296Crossref PubMed Scopus (1) Google Scholar). For ML algorithms on health administrative data, we represented data using count vectors of billing and diagnostic codes and the tf-idf algorithm and used various ML classifiers on these representations (e.g., random forest, logistic regression, shallow neural networks). The complete list of parameters and models can be found in Supplementary Tables S2–S6. We calculated sensitivity, specificity, positive predictive value (PPV), and negative predictive value with 95% confidence intervals (CIs) for each algorithm overall and stratified by age (children, adults). All models were evaluated using three-fold stratified cross validation. We considered a model better than another if its F1score (harmonic mean of PPV and sensitivity) was higher. The terms “eczema” or “dermatitis” were found in 863 of 2,036 (42.3%) child and 1,440 of 5,232 (27.5%) adult charts searched. AD was diagnosed according to the reference standard in 214 of 2,036 (10.5%) children and 116 of 5,232 (2.2%) adults, comparable with population prevalence estimates for AD in Canada (children: 11%, adults: 3%) (Bridgman et al., 2020Bridgman A.C. Fitzmaurice C. Dellavalle R.P. Karimkhani Aksut C. Grada A. Naghavi M. et al.Canadian burden of skin disease from 1990 to 2017: results from the Global Burden of Disease 2017 Study.J Cutan Med Surg. 2020; 24: 161-173Crossref PubMed Scopus (5) Google Scholar; Global Health Data Exchange, XXXGlobal Health Data ExchangeGBD results tool.http://ghdx.healthdata.org/gbd-results-toolDate: 2020Date accessed: July 24, 2020Google Scholar). Results for the best-performing rules-based and ML algorithms within EMRPC data were comparable, with the best ML algorithms performing slightly better. A simple classification rule within EMRPC searched for AD and atopic eczema (overall PPV = 81.7 [95% CI = 77.1–86.2]; overall sensitivity = 59.7 [95% CI = 58.3–61.1]) (Table 1 and Supplementary Tables S2–S4). The best-performing within-EMRPC ML model used logistic regression on human-defined features with a threshold probability set to 0.3 (overall PPV = 76.4 [95% CI = 73.6–79.1]; overall sensitivity = 67.3 [95% CI = 65.1–69.5]). These models all performed better in children than adults. All ML models trained on EMRPC free text rather than human-defined features performed poorly.Table 1PPV and Sensitivity of the Simple and Best-Performing Algorithms Within Each Approach to Identify AD in Ontario, Canada, Routinely Collected Health DataModelOverall Testing PopulationAdultsChildrenF1Score (95% CI)PPV (95% CI)Sensitivity (95% CI)PPV (95% CI)Sensitivity (95% CI)PPV (95% CI)Sensitivity (95% CI)Within EMRPC patient charts Rules-based algorithmsSimple algorithm: AD or eczema anywhere in chart68.9 (67.2–70.6)81.7 (77.1–86.2)59.7 (58.3–61.1)73.2 (63.8–82.7)59.5 (55.9–63.1)87.2 (85.2–89.1)59.9 (56.7–63.1)More complex algorithm1(AD or eczema anywhere in chart) or ([dermatitis or eczema anywhere in the chart] and [three visits associated with ICD-9 691 in 5 years] and [any visit associated with ICD-9 493 ever] and [prescription for topical anti-inflammatory therapy] and [asthma, rhinitis or hay fever anywhere in the chart]).69.7 (67.9–71.6)77.7 (73.9–81.4)63.3 (61.9–64.7)70.1 (58.0–82.2)62.2 (56.0–68.4)82.4 (80.4–84.4)63.7 (69.6–67.7) ML algorithmsLogistic regression on human-defined features2This model uses logistic regression with probability threshold being 0.3. The probability threshold is the minimum probability estimate required for a positive prediction. Variables in the regression model are sex; age group; presence of AD and/or eczema, atopic and/or atopy, eczema and/or dermatitis, asthma, rhinitis and/or hay fever, itch and/or scratch, or potential derivations and/or misspellings of those terms in the clinic note or cumulative patient profile; various combinations of ICD-9 codes for eczema (691), asthma (493), or rhinitis (477); and treatments (topical, phototherapy, systemic).71.5 (69.6–73.3)76.4 (73.6–79.1)67.3 (65.1–69.5)72.0 (68.5–75.5)60.3 (56.3–64.3)78.4 (75.4–81.5)71.2 (68.1–74.2)Logistic regression on free text3Logistic regression (class_weight = balanced). The hyperparameter class_weight represents the weights associated with each class, which affect how much impact each example from each class has. When class_weight is set to balanced, the weights are automatically adjusted to be inversely proportional to class frequency in the training set.41.0 (38.9–43.1)33.5 (29.6–37.3)54.2 (46.7–61.8)35.6 (33.2–38.0)11.1 (4.3–17.9)33.2 (29.0–37.4)77.5 (70.1–84.8)Within administrative health data One visit ever with ICD-9 69111.8 (10.9–12.7)6.3 (5.8–6.9)87.0 (84.6–89.4)3.2 (2.8–3.5)88 (84.7–91.4)14.5 (14.3–14.8)86.8 (84.8–88.9) Logistic regression3Logistic regression (class_weight = balanced). The hyperparameter class_weight represents the weights associated with each class, which affect how much impact each example from each class has. When class_weight is set to balanced, the weights are automatically adjusted to be inversely proportional to class frequency in the training set.19.6 (17.5–21.6)11.8 (10.6–13.0)58.5 (50.7–66.2)6.9 (6.5–7.2)26.0 (15.3–36.8)13.4 (11.7–15.1)77.1 (74.1–80.0)Abbreviations: AD, atopic dermatitis; CI, confidence interval; EMRPC, Electronic Medical Record Primary Care database; ICD-9, International Classification of Diseases 9; ML, machine learning; PPV, positive predictive value.Specificity and negative predictive values are presented in the Supplementary Tables.1 (AD or eczema anywhere in chart) or ([dermatitis or eczema anywhere in the chart] and [three visits associated with ICD-9 691 in 5 years] and [any visit associated with ICD-9 493 ever] and [prescription for topical anti-inflammatory therapy] and [asthma, rhinitis or hay fever anywhere in the chart]).2 This model uses logistic regression with probability threshold being 0.3. The probability threshold is the minimum probability estimate required for a positive prediction. Variables in the regression model are sex; age group; presence of AD and/or eczema, atopic and/or atopy, eczema and/or dermatitis, asthma, rhinitis and/or hay fever, itch and/or scratch, or potential derivations and/or misspellings of those terms in the clinic note or cumulative patient profile; various combinations of ICD-9 codes for eczema (691), asthma (493), or rhinitis (477); and treatments (topical, phototherapy, systemic).3 Logistic regression (class_weight = balanced). The hyperparameter class_weight represents the weights associated with each class, which affect how much impact each example from each class has. When class_weight is set to balanced, the weights are automatically adjusted to be inversely proportional to class frequency in the training set. Open table in a new tab Abbreviations: AD, atopic dermatitis; CI, confidence interval; EMRPC, Electronic Medical Record Primary Care database; ICD-9, International Classification of Diseases 9; ML, machine learning; PPV, positive predictive value. Specificity and negative predictive values are presented in the Supplementary Tables. Rules-based and ML algorithms on health administrative data had acceptable sensitivity but very low PPV, likely because of use of truncated International Classification of Diseases 9 code 691 for other eczemas and inflammatory skin conditions. Test properties of International Classification of Diseases 9 and 10 codes in other settings have been mixed, with poor performance in a United States quaternary care setting (Hsu et al., 2017Hsu D.Y. Dalal P. Sable K.A. Voruganti N. Nardone B. West D.P. et al.Validation of International Classification of Disease Ninth Revision codes for atopic dermatitis.Allergy. 2017; 72: 1091-1095Crossref PubMed Scopus (36) Google Scholar) but high PPV (95%) in Danish hospital dermatology clinics (Andersen et al., 2019Andersen Y.M.F. Egeberg A. Skov L. Thyssen J.P. Demographics, healthcare utilization and drug use in children and adults with atopic dermatitis in Denmark: a population-based cross-sectional study.J Eur Acad Dermatol Venereol. 2019; 33: 1133-1142Crossref PubMed Scopus (17) Google Scholar). When we assessed differences in within-EMRPC algorithm performance stratified by demographic factors, there were no clear patterns according to sex, rurality, or income quintile (Supplementary Tables S7 and S8). There were differences in the sensitivity of the logistic regression algorithm between different residential ethnic quintiles, but the small number of examples and large variances make interpreting these differences challenging. In summary, we identified simple rules-based and ML algorithms with adequate test properties to identify people with AD using EMRPC patient charts. Depending on the goals and resources of future studies, investigators may choose to implement either the simpler rules-based approach, with higher PPV, or logistic regression, with higher sensitivity. The overall PPV of 83% for the best-performing algorithm is comparable with those reported for algorithms to identify AD in United Kingdom electronic medical record data (86%) (Abuabara et al., 2017Abuabara K. Magyari A.M. Hoffstad O. Jabbar-Lopez Z.K. Smeeth L. Williams H.C. et al.Development and validation of an algorithm to accurately identify atopic eczema patients in primary care electronic health records from the UK.J Invest Dermatol. 2017; 137: 1655-1662Abstract Full Text Full Text PDF PubMed Scopus (27) Google Scholar) and British Columbia health administrative data (78%) (Zhang et al., 2020Zhang T. Lee T.K. Lui H. Dutz J. Dawes M. Lee A. et al.Health insurance claim- and prescription record-based algorithms as a population-based method for eczema ascertainment.J Eur Acad Dermatol Venereol. 2020; 34: e466-e468Crossref PubMed Scopus (1) Google Scholar). Algorithms using Ontario health administrative data had insufficient PPV to identify AD, but cohorts of people with and without AD created using algorithms within EMRPC patient charts can be linked with health administrative data to conduct reliable epidemiologic and health services research on AD in Ontario, Canada. Secure access to these data is governed by policies and procedures that are approved by the Information and Privacy Commissioner of Ontario, Canada. The data set from this study is held securely in encoded, deidentified form at ICES. Although data-sharing agreements prohibit ICES from making the data set publicly available, access may be granted to those who meet prespecified criteria for confidential access (www.ices.on.ca/DAS). Mohamed Abdalla: http://orcid.org/0000-0002-2776-6036 Branson Chen: http://orcid.org/0000-0002-4224-0753 Robin Santiago: http://orcid.org/0000-0001-9499-3768 Jacqueline Young: http://orcid.org/0000-0003-0994-1934 Lihi Eder: http://orcid.org/0000-0002-1473-1715 An-Wen Chan: http://orcid.org/0000-0002-4498-3382 Elena Pope: http://orcid.org/0000-0002-2136-5661 Karen Tu: http://orcid.org/0000-0003-0883-4934 Liisa Jaakkimainen: http://orcid.org/0000-0002-3203-0007 Aaron M. Drucker: http://orcid.org/0000-0002-7388-9475 RS was employed by ICES when his contributions were made. LE reports receiving educational grants from and serves on the advisory board for Novartis, Eli Lily, Jensen, AbbVie, and Pfizer. AMD has been a paid consultant for Sanofi, RTI Health Solutions, Eczema Society of Canada, and Canadian Agency for Drugs and Technology in Health. He has received honoraria from CME Outfitters. His institution has received educational grants from Sanofi and research grants from Sanofi and Regeneron. The remaining authors state no conflict of interest. We thank Li Bai for contributing to the statistical analysis. This work was funded by research grants from the Canadian Dermatology Foundation , the Physicians Services Incorporated Foundation, and the Canadian Institutes of Health Research . MA is supported by a Vanier Canada Graduate Scholarship. This study was supported by ICES, which is funded by an annual grant from the Ontario Ministry of Health and Long-Term Care . Parts of this material are based on data and/or information compiled and provided by the Canadian Institute for Health Information. The analyses, conclusions, opinions, and statements expressed here are solely those of the authors and do not reflect those of the funding or data sources; no endorsement is intended or should be inferred. The analyses, conclusions, opinions, and statements expressed in the material are those of the author(s) and not necessarily those of Canadian Institute for Health Information. Conceptualization: AMD, KT, LJ, AWC, LE, EP, MA; Data Curation: JY, BC, RS, MA; Formal Analysis: BC, RS, MA; Funding Acquisition: AMD, KT, LJ, AWC, LE; Investigation: AMD, LJ, MA, JY, BC, RS; Methodology: AMD, KT, LJ, MA; Project Administration: AMD; Resources: AMD, LJ; Software: LJ, MA; Supervision: AMD, LJ, KT; Validation: BC, JY, MA, RS; Writing - Original Draft Preparation: AMD, MA; Writing - Review and Editing: MA, BC, RS, JY, LE, AWC, EP, KT, LJ, AMD Download .pdf (.41 MB) Help with pdf files Supplementary Material

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.023
metaresearch head score (Gemma)0.187
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.716
Threshold uncertainty score0.565

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0230.187
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.004
Bibliometrics0.0100.010
Science and technology studies0.0010.001
Scholarly communication0.0050.003
Open science0.0030.003
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.042
GPT teacher head0.334
Teacher spread0.292 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2021
Admission routes3
Has abstractyes

Explore more

Same venueJournal of Investigative DermatologySame topicDermatology and Skin DiseasesFrench-language works237,207