Identification of undiagnosed diabetes and quality of diabetes care in the United States: cross-sectional study of 11.5 million primary care electronic records
Bibliographic record
Abstract
BACKGROUND: Electronic diabetes registers promote structured care and enable identification of undiagnosed diabetes, but they require consistent coding of the diagnosis in electronic medical records. We investigated the potential of electronic medical records to identify undiagnosed diabetes and to support diabetes management in a large primary care population in the United States. METHODS: We conducted a cross-sectional study and retrospective observational cohort analysis of primary care electronic medical records from a nationally representative US database (GE Centricity). We tested the feasibility of identifying patients with undiagnosed diabetes by applying simple algorithms to the electronic medical record data. We compared the quality of care provided to patients in the United States who had diabetes (coded and uncoded) for at least 15 months with the quality of care provided in England using a set of 16 indicators. RESULTS: We included 11 540 454 electronic medical records from more than 9000 primary care clinics across the United States. Of the 1 110 398 records indicating diagnosed diabetes, only 61.9% contained a diagnostic code. Of the 10 430 056 records for nondiabetic patients, 0.4% (n = 40 359) had at least 2 abnormal fasting or random blood glucose values, and 0.2% (n = 23 261) of the remaining records had at least 1 documented glycated hemoglobin (HbA1c) value of 6.5% or higher. Among the 622 260 patients for whom information on quality-of-care indicators was available, those with a coded diagnosis of diabetes had a significantly higher level of quality of care than those with uncoded diabetes (p < 0.01); however, the quality of care was generally lower than that indicated in England. INTERPRETATION: We were able to identify a substantial number of patients with uncoded diabetes and probable undiagnosed diabetes using simple algorithms applied to the primary care electronic records. Electronic coding of the diagnosis was associated with improved quality of care. Electronic diabetes registers are underused in US primary care and provide opportunities to facilitate the systematic, structured approach that is established in England.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".