Using Linked Administrative Health Databases for An Obesity Case Definition
Bibliographic record
Abstract
IntroductionCapture of obesity using administrative health data is poor, with many cases being under coded within the data. Linking multiple health data sources may improve case ascertainment and facilitate the use of administrative health data for obesity research and surveillance.
 Objectives and ApproachThis research aims to determine if using individual-level linked data from multiple sources can improve case ascertainment for obesity in administrative health data. Data from between April 1, 2001 and March 31, 2015 were obtained from the Manitoba Population Data Repository. Eighteen obesity case definitions were developed with different observation times and combinations of diagnosis, procedure, and prescription codes from physician billing claims, hospitalization abstracts, and prescription drug records. Body mass index (BMI) records from primary care data and the Bone Mineral Density (BMD) registry were used for validation. Sensitivity, specificity, and Cohen’s kappa were calculated.
 ResultsIndividuals with a higher BMI class had more physician visits and were more likely to have comorbidities and obese codes in the administrative health data. A higher BMI class was associated with being in a lower income quintile and the age group 40-59. Overall, the case definitions for obesity had high specificity (0.98-0.99) and low sensitivity (0.005-0.19) when validated using primary care data. Case definitions with obesity codes from multiple databases 3 year prior to and including the index date had the highest sensitivity (0.06-0.19) and kappa (0.04-0.23). Results with the BMD data were similar (specificity: 0.97-0.99; sensitivity: 0.007-0.21). Stratified analyses found agreement measures improved slightly for females, those who had chronic conditions and a later index year, and the age group 40-59.
 Conclusion / ImplicationsWhen using multiple databases to build a case definition for obesity, sensitivity improves but remains low. Individuals with other chronic conditions and a higher BMI class were more likely to be accurately classified as an obese case.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".