Multitrait genetic-phenotype associations to connect disease variants and biological mechanisms
Bibliographic record
Abstract
Abstract Background Genome-wide association studies (GWAS) uncovered a wealth of associations between common variants and human phenotypes. These results, widely shared across the scientific community as summary statistics, fostered a flurry of secondary analysis: heritability and genetic correlation assessment, pleiotropy characterization and multitrait association test. Amongst these secondary analyses, a rising new field is the decomposition of multitrait genetic effects into distinct profiles of pleiotropy. Results We conducted an integrative analysis of GWAS summary statistics from 36 phenotypes to decipher multitrait genetic architecture and its link to biological mechanisms. We started by benchmarking multitrait association tests on a large panel of phenotype sets and established the Omnibus test as the most powerful in practice. We detected 322 new associations that were not previously reported by univariate screening. Using independent significant associations, we investigated the breakdown of genetic association into clusters of variants harboring similar multitrait association profile. Focusing on two subsets of immunity and metabolism phenotypes, we then demonstrate how SNPs within clusters can be mapped to biological pathways and disease mechanisms, providing a putative insight for numerous SNPs with unknown biological function. Finally, for the metabolism set, we investigate the link between gene cluster assignment and success of drug targets in random control trials. We report additional uninvestigated drug targets classified by clusters. Conclusions Multitrait genetic signals can be decomposed into distinct pleiotropy profiles that reveal consistent with pathways databases and random control trials. We propose this method for the mapping of unannotated SNPs to putative pathways.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.022 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".