Incorporating Alternative Polygenic Risk Scores into the BOADICEA Breast Cancer Risk Prediction Model
Bibliographic record
Abstract
BACKGROUND: The multifactorial risk prediction model BOADICEA enables identification of women at higher or lower risk of developing breast cancer. BOADICEA models genetic susceptibility in terms of the effects of rare variants in breast cancer susceptibility genes and a polygenic component, decomposed into an unmeasured and a measured component - the polygenic risk score (PRS). The current version was developed using a 313 SNP PRS. Here, we evaluated approaches to incorporating this PRS and alternative PRS in BOADICEA. METHODS: The mean, SD, and proportion of the overall polygenic component explained by the PRS (α2) need to be estimated. $\alpha $ was estimated using logistic regression, where the age-specific log-OR is constrained to be a function of the age-dependent polygenic relative risk in BOADICEA; and using a retrospective likelihood (RL) approach that models, in addition, the unmeasured polygenic component. RESULTS: Parameters were computed for 11 PRS, including 6 variations of the 313 SNP PRS used in clinical trials and implementation studies. The logistic regression approach underestimates $\alpha $, as compared with the RL estimates. The RL $\alpha $ estimates were very close to those obtained by assuming proportionality to the OR per 1 SD, with the constant of proportionality estimated using the 313 SNP PRS. Small variations in the SNPs included in the PRS can lead to large differences in the mean. CONCLUSIONS: BOADICEA can be readily adapted to different PRS in a manner that maintains consistency of the model. IMPACT: : The methods described facilitate comprehensive breast cancer risk assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".