Multi-Ancestry Proteo-Genomic Association Study of eGFR
Bibliographic record
Abstract
Background: Reduced eGFR impacts the concentration of proteins circulating in the plasma. Biomarkers may be released during injury without being harmful themselves. Thus, biomarker studies of CKD are prone to confounding by reverse causation. The concentration of plasma protein biomarkers is also influenced by genetic variation. As genotypes are static from birth, life-long differences in genetically predicted protein concentration can be used to identify potentially causal pathogenic or protective proteins while minimizing confounding and reverse causality. Jointly considering the impact of multiple variants on life-long protein concentration can discover novel associations in addition to variants identified in genome-wide association studies (GWAS). In a proteogenomic association study, we sought to test the impact of genetically predicted variation in 1,161 plasma proteins on eGFR. Methods: We searched for cis protein quantitative trait loci (pQTL) genetic variants associated with the concentration of 1,161 plasma proteins in a multi-ancestry sample of 10,753 participants from the Prospective Urban and Rural Epidemiological (PURE) study. Using two-sample Mendelian randomization, we tested if pQTL variants were also associated with eGFR and kidney traits in >1 million participants of published GWAS. We also examined colocalization of pQTL signals and eGFR GWAS results and the phenomewide impact of genetically altered concentration of the identified proteins. Results: 419 pQTL instruments were constructed in PURE including 4665 genetic variants. In GWAS data, genetically altered concentration of 27 protein biomarkers was associated with eGFR (P < 9.5 x 10-5). UMOD was the strongest signal, a positive control of the analysis. Six of the significant biomarkers were previously identified in GWAS; 12 were in identified loci but the causal gene under the GWAS peak was unknown; and 9 loci were unidentified in GWAS. Novel biomarker associations with eGFR include interesting biological candidates such as inhibin beta chain C, a subunit of activins, and proteinase-3, the antigen in PR3-ANCA vasculitis. Conclusions: Using a proteo-genomic association study, 27 biomarkers whose genetically predicted concentration were causally associated with changes in kidney function and risk of kidney disease including interesting biological candidates. Funding: Commercial Support - Bayer
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".