Exploring novel diabetes surveillance methods: a comparison of administrative, laboratory and pharmacy data case definitions using THIN
Bibliographic record
Abstract
Background: The objective of this study was to identify patients with diabetes in a comprehensive primary care electronic medical records database using a number of different case definitions (clinical, pharmacy, laboratory definitions and a combination thereof) and understand the differences in patient populations being captured by each definition. Methods: Data for this population-based retrospective cohort study was obtained from The Health Information Network (THIN). THIN is a longitudinal, primary care medical records database of over 9 million patients in UK. Primary outcome was a diagnosis of diabetes, defined by the presence of a diabetes read code, or an abnormal laboratory result, or a prescription for an Oral Anti-diabetic drug or insulin. A 2-year washout period was applied prior to the index of diabetes to avoid inclusion of prevalent cases for each case definition. Results: This study demonstrated that different case definitions of diabetes identify different sub-populations of patients. When the cohorts were observed based on any measure of central tendency, each of the cohorts were reasonably comparable to each other. However, the distribution of each of the cohorts when grouped by age categories and sex, reveal differences. For example, using pharmacy case definition results in a bimodal distribution among women, one between 1-19 year and 35-39 age categories, and then again between 60-64 and 85 years-however, the histogram becomes more normally distributed when metformin was removed from the case definition. Conclusion: Our results suggest that clinical, pharmacy, laboratory case definitions identify different sub-populations and using multiple case definitions is likely required to optimally identify the entire diabetes population within THIN. Our study also suggests that age and sex of patients may affect the indexing of diabetes in THIN and is critical to better understand these variations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".