MétaCan
Menu
Back to cohort
Record W2948406734 · doi:10.2337/db19-1672-p

1672-P: Identifying Type 1 Diabetes in Electronic Medical Records Data in Ontario: A Validation Study

2019· article· en· W2948406734 on OpenAlexaboutno aff
Alanna Weisman, Jacqueline Young, Matthew Kumar, Peter C. Austin, Karen Tu, Liisa Jaakkimainen, Lorraine L. Lipscombe, Gillian L. Booth

Bibliographic record

VenueDiabetes · 2019
Typearticle
Languageen
FieldMedicine
TopicDiabetes Management and Research
Canadian institutionsnot available
Fundersnot available
KeywordsMedicineMedical recordCartMetforminCohortType 2 diabetesDiagnosis codeDiabetes mellitusPediatricsCohort studyInternal medicineRetrospective cohort studyElectronic medical recordEmergency medicineInsulinPopulationEndocrinology

Abstract

fetched live from OpenAlex

Electronic Medical Record (EMR) data are an efficient method for constructing large type 1 diabetes (T1D) cohorts but are limited by availability of accurate diagnosis codes. We sought to derive an algorithm for identifying T1D using Electronic Medical Records - Primary Care (EMRPC) also known as EMRALD, a database including records from >300,000 primary care patients in Ontario, Canada. A 25% random sample of adults with diabetes and at least 1 year of records in EMRPC prior to September 30, 2015 formed the reference cohort. Charts of potential T1D cases (those on insulin±metformin) were abstracted; all others were classified as non-T1D. Algorithms were derived using variables from free-text searches (diagnosis terms in the Cumulative Patient Profile (CPP)) and standardized fields (medications, age, BMI) using 1)a priori combinations of variables; 2)classification and regression tree analysis (CART). Algorithms were evaluated for sensitivity, specificity, positive and negative predictive values (PPV and NPV) with adjustment for optimism for CART. Of 21,547 eligible patients with diabetes, the reference cohort was 5407: 4968 non-abstracted non-T1D, 240 abstracted non-T1D and 199 T1D. The prevalence of T1D was 3.7%. The optimal algorithm using a priori variable combinations was T1D diagnosis in the CPP + rapid-acting insulin (sensitivity 69.3% (95% CI 62.4-75.7%), specificity 99.8% (99.6-99.9), PPV 92.6% (87.2-96.3), NPV 98.8% (98.5-99.1)). CART partitioned on T1D diagnosis in the CPP, any insulin, and non-metformin oral hypoglycemic medications and improved performance with optimism-adjusted sensitivity of 77% (72.0-83.9), specificity 99.8% (99.8-100.0), PPV 96.8% (94.6-99.6), and NPV 99.1% (98.9-99.4). Simple algorithms using EMR variables yielded good diagnostic performance for identification of T1D and CART further improved performance. Pending additional validation these algorithms can be applied to study large T1D cohorts in EMR databases. Disclosure A. Weisman: None. J. Young: None. M. Kumar: None. P. Austin: None. K. Tu: None. L. Jaakkimainen: None. L. Lipscombe: None. G. Booth: None. Funding Diabetes Action Canada

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.023
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.103
Threshold uncertainty score0.208

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.023
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0010.002
Science and technology studies0.0020.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.049
GPT teacher head0.341
Teacher spread0.292 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueDiabetesSame topicDiabetes Management and ResearchFrench-language works237,207