Phenome-wide analysis of <i>APOL1</i> risk variants reveals associations between one combination of haplotypes and multiple disease phenotypes in addition to chronic kidney disease
Bibliographic record
Abstract
Abstract Background Infectious diseases are a major driving force of natural selection. One human gene associated with strong evolutionary selection is APOL1 . Two APOL1 variants, G1 and G2, emerged in sub-Saharan Africa in the last 10,000 years, possibly due to protection from the fatal African sleeping sickness, analogous to Plasmodium -driven selection of the sickle-cell trait. As homozygosity for the HbS allele causes sickle cell anaemia, homozygosity for the APOL1 G1 and G2 variants has also been associated with chronic kidney disease (CKD) and other kidney-related conditions. What is not known is the extend of non-kidney-related disorders and if there are clusters of diseases associated with individual APOL1 genotypes. Methods Using principal component analysis, we identified a cohort of 10,179 UK Biobank participants with recent African ancestry. We conducted a phenome-wide association test between all combinations of APOL1 G1 and G2 genotypes and conditions identified with International Classification of Disease phenotypes using Firth’s bias-reduced logistic regression and a false discovery rate to correct for multiple testing. We further examined associations with chronic kidney disease indicators: estimated glomerular filtration rate (eGFR) and urinary albumin:creatinine (uACR). Results The phenome-wide screen revealed 74 (mostly deleterious) potential associations with hospitalisation for a range of conditions. G1/G2 compound heterozygotes were specifically associated with hospitalisation in 64 (86.5%) of these conditions, with an over-representation of infectious diseases (including COVID-19) and endocrine, nutritional, and metabolic diseases. The analysis also revealed complexities in the relationship between APOL1 and CKD that are not evident when the risk variants are grouped together: high uACR was associated specifically with G1 homozygosity; low eGFR with G2 homozygosity and G1/G2 compound heterozygosity; progression to end stage kidney disease was associated with G1/G2 compound heterozygosity. Conclusions Among 9,594 participants, stratifying individual APOL1 risk variant genotypes had a differential effect on associations with both kidney and non-kidney phenotypes. The compound heterozygous G1/G2 genotype was distinguished as uniquely deleterious in its association with a range of ICD-10 phenotypes. The epistatic nature of the G1/G2 interaction means that such associations may go undetected in a standard genome-wide association study. These observations have the potential to significantly impact the way that health risks are understood, particularly in populations where APOL1 G1 and G2 are common such as in sub-Saharan Africa and its diaspora.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".