Estimating Disease Heritability from Electronic Healthcare Records: A Proof-of-Concept Study.
Bibliographic record
Abstract
ObjectiveA family history of a chronic disease often predicts disease risk, with predictive value determined by heritability, the proportion of variation in risk explained by inherited genetic factors. Our objective was to assess the validity of disease heritability estimates from electronic healthcare records (EHRs) that capture family relationships and disease diagnoses. ApproachA population-based investigation was conducted using healthcare records from Manitoba, Canada for 1970 to 2021. We constructed family relationships for up to four generations using health insurance registration information containing unique family and individual identifiers. Health histories for family members were created using diagnosis codes in hospital and physician visit records. Linear mixed-effects models were used to estimate heritability (h) for 130 chronic health conditions using open-source Clinical Classifications Software that defines clinically-meaningful disease categories. Comparisons between EHR-derived estimates and genetically-derived estimates from published studies were used to assess validity of the methodology. ResultsHealth insurance registration data were used to construct 10,000 families that included 116,879 individuals. Median family size was 9 (interquartile range: 8). Median observation time was 39.6 years (interquartile range: 25.7). Males comprised half (51.0%) of family members. A total of 272,114 familial relationships were identified; slightly more than half (53%) were first degree (i.e., child and parent) relationships. One-third (33.2%) of families were comprised of four generations; only 15.3% were comprised of two generations. Heritability estimates were consistent with published genetically-derived estimates for several conditions, including diabetes (EHR h = 0.29 vs. 0.22), anemia (EHR h = 0.21 vs. 0.20), and asthma (EHR h = 0.34 vs. 0.33). However, inconsistencies were identified for pancreatic disorders, gastrointestinal conditions, some mental health conditions, and heart disease. ConclusionEHRs provide a promising approach to explore heritability of selected health conditions in large, diverse populations. Inconsistencies between EHR-derived and genetically-derived estimates are indicative of the limitations of diagnoses recorded for administrative purposes. Future research will explore sex-specific heritability estimates and effects of change in disease diagnosis coding over time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.106 | 0.211 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".