Enhancing the BOADICEA cancer risk prediction model to incorporate new data on <i>RAD51C</i> , <i>RAD51D</i> , <i>BARD1</i> updates to tumour pathology and cancer incidence
Bibliographic record
Abstract
Background BOADICEA (Breast and Ovarian Analysis of Disease Incidence and Carrier Estimation Algorithm) for breast cancer and the epithelial tubo-ovarian cancer (EOC) models included in the CanRisk tool ( www.canrisk.org ) provide future cancer risks based on pathogenic variants in cancer-susceptibility genes, polygenic risk scores, breast density, questionnaire-based risk factors and family history. Here, we extend the models to include the effects of pathogenic variants in recently established breast cancer and EOC susceptibility genes, up-to-date age-specific pathology distributions and continuous risk factors. Methods BOADICEA was extended to further incorporate the associations of pathogenic variants in BARD1 , RAD51C and RAD51D with breast cancer risk. The EOC model was extended to include the association of PALB2 pathogenic variants with EOC risk. Age-specific distributions of oestrogen-receptor-negative and triple-negative breast cancer status for pathogenic variant carriers in these genes and CHEK2 and ATM were also incorporated. A novel method to include continuous risk factors was developed, exemplified by including adult height as continuous. Results BARD1 , RAD51C and RAD51D explain 0.31% of the breast cancer polygenic variance. When incorporated into the multifactorial model, 34%–44% of these carriers would be reclassified to the near-population and 15%–22% to the high-risk categories based on the UK National Institute for Health and Care Excellence guidelines. Under the EOC multifactorial model, 62%, 35% and 3% of PALB2 carriers have lifetime EOC risks of <5%, 5%–10% and >10%, respectively. Including height as continuous, increased the breast cancer relative risk variance from 0.002 to 0.010. Conclusions These extensions will allow for better personalised risks for BARD1 , RAD51C , RAD51D and PALB2 pathogenic variant carriers and more informed choices on screening, prevention, risk factor modification or other risk-reducing options.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".