Missing heritability or need for reality check of clinical utility in genomic testing?
Bibliographic record
Abstract
Missing heritability is one of the most hotly debated subject in human genetics, based on the fact that early genome-wide association studies (GWAS), although leading to the identification of over 1000 genetic variants, associated with over 160 common human diseases and traits; these genetic variants have provided modest contribution to ‘heritability’ of the traits. With recent larger GWAS, including those from mega-consortia involving hundreds of thousands of participants, the proportion of heritability of the traits explained by genetic variants has grown to 20–30% in some cases and even to greater than 50% in a few, but for most traits, the majority of heritability remains unexplained [1]. Although many reasons were offered, including lack of consideration of gene × environment interaction, additional layers of uncovered complexity of genome and its function, impact of rare alleles, selective parental origin of risk allele, trans-generational phenotype modifications, unrecognized influence of copy number variances (CNV) and others, the problem is considered largely unresolved [2]. We would like to suggest that part of it is the story of a glass being half full or half empty in the eye of the beholder. What are our expectations? Genetic testing cannot resolve more than there is heritability, which in many cases is less than 50% for common complex polygenic disorders. Thus, hypertension heritability is estimated to be in the order of 30–45%, and of 49% for albuminuria and glomerular filtration rate (GFR) in the Hypertension Genetic Epidemiology Network (HyperGEN) families enriched for multiple siblings with hypertension [3]. Furthermore, Zuk et al.[1] suggested that a significant proportion of missing heritability could come from an overestimation of the heritability itself that considers strictly additive genetic model and not gene interaction, thereby creating ‘phantom heritability’. Furthermore, we tend to forget that heritability estimates are environment-dependent and that until we characterize the environment more rigorously, we have to consider these heritability estimates, and attempts to explain the underlying contributions of genes, as speculative, although progress is being made [4]. A similar debate could be made around ‘missing environment’. For instance, when all traditional cardiovascular risk factors were considered, even in conjunction with lifestyle behaviors, the percentage of fatal and nonfatal cardiovascular events explained by these factors was limited to 38 and 67%, respectively, in the PRIME study [3,5]. Furthermore, it is frequently cited that determination of family history-attributable risk has the same value as genetic factors without the necessity of costly genotyping. Family history is indeed a powerful clinical information that medicine uses since its scientific conception. Its capacity to attribute risk is significant at the level of patients’ stratification according to the level of risk. Thus, when a family history is included in a risk prediction model, such as in Framingham score for hypertension risk, it is indeed a predictor, since all genetic information from both parents is included [6]. However, its impact as an individual predictor, within a sib-ship, is a coin toasting. Similarly, when genetic factors of cardiovascular diseases (CVDs) are compared to the so-called ‘nongenetic factors’ or ‘traditional risk factors’, which typically include BMI, presence of diabetes, dyslipidemia, hypertension and even atrial fibrillation, albuminuria and other validated risk factors, we have to remind ourselves that these factors have well demonstrated genomic determinants, contributing at least in part to their effect size. Whereas such comparisons are justified and needed, a critical appreciation of its ‘nongenetic’ content is also required. This issue of Journal of Hypertension includes an elegant study from Malmö longitudinal observation, evaluating the cardiovascular consequences of polygenic character of hypertension [7]. Fava et al.[7] studied genetic polymorphisms derived from one of the largest GWAS meta-analysis of 120 000 participants that led to the identification of 29 significant single-nucleotide polymorphisms (SNPs) and demonstrated their impact on cardiovascular consequences on incident and prevalent stroke and coronary artery disease, but not on renal impairment [8]. Fava et al. confirmed the highly significant association of 24 of these SNPs with both SBP (β = 2.83 mmHg, P = 8.6 × 10−55) and DBP (β = 1.59 mmHg, P = 1.06 × 10−58), as well as with hypertension [odds ratio (OR) 1.32, P = 1.34 × 10−40]. Genetic risk significantly increased the risk of stroke, coronary artery disease, and cardiovascular mortality, but not of total mortality. The authors underlined and based their final conclusion on the fact that after adjustment for ‘traditional risk factors’, the genetic score ‘remained significantly associated only with CVDs (in terms of stroke and coronary artery disease) [hazard ratio (1.15), 95% confidence interval (CI) 1.06–1.24]. Yet, we have to mention that the traditional risk factors used for adjustment included age, sex, age2, age × sex, BMI, hypertension, diabetes, smoking, and use of antilipemic drugs. The strength of the study is that it is population-based. Perhaps the most significant finding is reported at the end of the ‘Discussion’ section: 1 SD of genetic score increases SBP by 2.8 mmHg and corresponds to 8% increase of CVD, whereas traditional risk factors increased it by 1 mmHg and resulted in a 1.5% increase of risk of CVD. This in itself strongly suggests the potential clinical utility of the data in outcome prediction. Why then the authors concluded in their Perspective that the ‘clinical importance for risk prediction among middle-aged individuals appears to be limited’. We believe that the half empty glass is the fact, demonstrated here that for a middle-aged man with high BMI, hypertension and diabetes, who smokes and has high low-density lipoprotein (LDL), genotyping is an unnecessary step to conclude that that person is at high CVD risk, but we also believe that the half full glass is for those who, in spite of highly heritable risk received from either parent before hypertension, obesity and dyslipidemia be detected, could benefit from prevention. The strength of genetic markers is in their presence from birth, allowing developing preventive and therapeutic measures at prime time, contrasting with current biomarkers of processes already initiated. This proposed path will need prospective confirmation of its clinical utility, novel strategies, and development of pro-active medicine, contrasting with current reactive mode. ACKNOWLEDGEMENTS The author would like to thank Professor Johanne Tremblay, PhD, FAHA, FCAHS from CHUM Research Center, Montreal, for fruitful discussion. Conflicts of interest There are no conflicts of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.259 | 0.605 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.017 |
| Scholarly communication | 0.008 | 0.018 |
| Open science | 0.006 | 0.007 |
| Research integrity | 0.008 | 0.013 |
| Insufficient payload (model declined to judge) | 0.009 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".