MétaCan
Menu
Back to cohort

Validation and Comparison of Algorithms to Identify Adult-Onset Inflammatory Bowel Disease Patients From Within Health Administrative Data

2012· article· en· W2330921232 on OpenAlexaffabout
Eric I. Benchimol, Astrid Guttmann, David Mack, Geoffrey C. Nguyen, Alan J. Forster, James C. Gregor, John K. Marshall, Douglas G. Manuel

Bibliographic record

VenueInflammatory Bowel Diseases · 2012
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicInflammatory Bowel Disease
Canadian institutionsWestern UniversityMcMaster UniversityUniversity of TorontoUniversity of Ottawa
Fundersnot available
KeywordsMedicineInflammatory bowel diseaseCohortDiseaseIncidence (geometry)Diagnosis codePopulationAlgorithmCrohn's diseaseHealth careCohort studyPediatricsInternal medicineComputer scienceEnvironmental health

Abstract

fetched live from OpenAlex

Health administrative databases can be used to track disease incidence, outcomes and care quality. Case validation is necessary to ensure accurate disease ascertainment using these databases. Disease-specific codes and algorithms (combinations of disease-specific health services codes) have been validated in various regions to identify adult-onset IBD. We previously validated an algorithm to identify pediatric-onset IBD and create the Ontario Crohn's and Colitis Cohort (OCCC), a population-based surveillance cohort. In this study, we aimed to validate adult-onset IBD identification algorithms to expand the OCCC to adults. Two cohorts of patients were used to validate algorithms: 1) Ottawa Hospital Algorithm Development Cohort: An electronic database (Ottawa Hospital Data Warehouse) of outpatient visits, hospitalizations, pathology and radiology reports was searched. The charts of patients with either a diagnostic ICD code for Crohn's or UC, or a referral to IBD, Crohn's or UC in a keyword search of either a histology or radiology report were extracted (n = 5847). Patients >18 years seen at the Ottawa Hospital from fiscal years (FY) 2002-2005 were classified as incident IBD (n = 554), prevalent IBD (n = 1193) or non-IBD (n = 3330). 2) Ontario Algorithm Validation Cohort: The charts from 5 community practices (family medicine and gastroenterology) and 3 tertiary care centers across Ontario, comprising the practices of >30 physicians, were extracted to classify IBD patients diagnosed FY2001-2005 (n = 464) and non-IBD patients (n = 1051). We linked to health administrative databases from FY1991-2010 and compared the accuracy of various algorithms with various lengths of time to qualify. In addition, their latest diagnosis (Crohn's or UC) was determined from charts to validate an algorithm to distinguish patients with Crohn's from those with UC. Over 5000 algorithms were tested. Table 1 details the diagnostic accuracies of published algorithms. One health care contact was not adequate to identify patients with IBD. The Manitoba algorithm attained the lowest false-positive rate, while maintaining sensitivity. The algorithms functioned variably by age, with diagnostic accuracy lower in patients with onset >65 years (e.g. Ottawa Hospital Algorithm Development Cohort for the Manitoba algorithm: Sens 59.3%, Spec 98.2%, PPV 58.2%, NPV 98.3%). Having 5 of the last 9 physician billing codes for Crohn's or UC accurately distinguished IBD subtypes (accuracy 91.1%), while having 4 of the last 7 codes (accuracy 90.7%) and 5 of the last 8 codes (accuracy 90.2%) also functioned adequately. Patients with adult-onset IBD can be accurately identified from within health administrative data. The previously validated algorithms from Manitoba were the most accurate. Further work will explore other algorithms, particularly in patient subgroups such as the elderly. The OCCC will be expanded to include adult patients, creating a large, population-based surveillance cohort of all patients with IBD in Ontario, Canada.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.051
metaresearch head score (Gemma)0.099
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.054
Threshold uncertainty score0.270

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0510.099
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.003
Science and technology studies0.0010.001
Scholarly communication0.0030.001
Open science0.0020.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.030
GPT teacher head0.335
Teacher spread0.305 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2012
Admission routes2
Has abstractyes

Explore more

Same venueInflammatory Bowel DiseasesSame topicInflammatory Bowel DiseaseFrench-language works237,207