Polygenic risk and rare variant gene clustering enhance cancer risk stratification for breast and prostate cancers
Bibliographic record
Abstract
Polygenic risk score (PRS) and rare monogenic variant screening are valuable tools for predicting cancer risk and identifying individuals at high risk. Integrating both common and rare genetic variants is crucial for accurate risk assessment. However, estimating the impacts of rare variants on cancer and combining them with PRS remains challenging. Here, we analyze 454,711 exome sequencing and 487,409 array UK Biobank samples, focusing on breast and prostate cancers. We introduce an expanded PRS (EPRS) approach, yielding a systematic model for more effective risk stratification. By prioritizing and clustering genes with cancer-specific rare variants based on odds ratios and population-attributable fraction, we refine risk stratification by combining both monogenic and polygenic effects. Individuals in high-PRS groups with rare high-impact gene variants show up to 15- and 22-fold higher risk for breast and prostate cancers, respectively, compared to those in the intermediate-PRS groups without rare variants. Combined risk profiles vary across distinct rare variant clusters within the same PRS group for both cancers. Our EPRS approach enhances risk stratification for breast and prostate cancers, offering important insights for future research and potential applications to other cancer types. An expanded PRS (EPRS) approach combining polygenic risk scores and rare variant clustering enhances cancer risk stratification for breast and prostate cancers. High-PRS groups with rare high-impact gene variants have up to 15- and 22-fold higher risk for breast and prostate cancers, respectively, compared to intermediate-PRS groups without rare variants.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".