Overlap of high-risk individuals across family history, genetic & non-genetic breast cancer risk models: Analysis of 180,398 women from European & Asian ancestries
Bibliographic record
Abstract
ABSTRACT Background Breast cancer is multifactorial. Focusing on limited risk factors may miss high-risk individuals. Methods We assessed the performance and overlap of various risk factors in identifying high-risk individuals for invasive breast cancer (BrCa) and ductal carcinoma in situ (DCIS) in 161,849 European-ancestry and 18,549 Asian-ancestry women. Discriminatory ability was evaluated using the area under the receiver operating characteristic curve (AUC). High-risk criteria included: 5-year absolute risk ≥1·66% by the Gail model [GAIL binary ]; first-degree family history of breast cancer [FH binary ]; 5-year absolute risk ≥1·66% by a 313-variants polygenic risk score [PRS binary ]; and carriers of pathogenic variants in breast cancer predisposition genes [PTV binary ]. Findings The 5-year absolute risk by PRS outperformed the Gail model in predicting BrCa (Europeans vs controls : AUC PRS =0·635 [0·632-0·638] vs AUC Gail =0·492 [0·489-0·495]; Asians vs controls : AUC PRS =0·564 [0·556-0·573] vs AUC Gail =0·506 [0·497-0·514]). PRS binary and GAIL binary identified more high-risk European than Asia individuals. High-risk proportions were higher among BrCa (16-26%) and DCIS (20-33%) compared to controls (9-15%) among young Europeans and all Asians. Fewer than 7% of BrCa, 10% of DCIS, and 3% of controls were classified as high-risk by multiple risk classifiers. Overlap between PRS binary and PTV binary was minimal (<0·65% Europeans, <0·15% Asians) compared to the proportion at high risk using PTV binary alone (Europeans: 4·6%, Asians: 4·4%) and PRS binary alone (Europeans: 13·9%, Asians: 8·5%). PRS binary and FH binary uniquely identified 5-6% and 9-11% of young BrCa, respectively. Interpretation The incomplete overlap between high-risk individuals identified by PRS binary , GAIL binary , FH binary, and PTV binary highlights the need for a comprehensive approach to breast cancer risk prediction. SIGNIFICANCE This study shows that different ways of predicting breast cancer risk do not always flag the same people, suggesting that combining multiple risk factors could improve early detection and screening.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.012 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".