Prediction of breast cancer risk based on common genetic variants in women of East Asian ancestry
Bibliographic record
Abstract
BACKGROUND: Approximately 100 common breast cancer susceptibility alleles have been identified in genome-wide association studies (GWAS). The utility of these variants in breast cancer risk prediction models has not been evaluated adequately in women of Asian ancestry. METHODS: We evaluated 88 breast cancer risk variants that were identified previously by GWAS in 11,760 cases and 11,612 controls of Asian ancestry. SNPs confirmed to be associated with breast cancer risk in Asian women were used to construct a polygenic risk score (PRS). The relative and absolute risks of breast cancer by the PRS percentiles were estimated based on the PRS distribution, and were used to stratify women into different levels of breast cancer risk. RESULTS: We confirmed significant associations with breast cancer risk for SNPs in 44 of the 78 previously reported loci at P < 0.05. Compared with women in the middle quintile of the PRS, women in the top 1% group had a 2.70-fold elevated risk of breast cancer (95% CI: 2.15-3.40). The risk prediction model with the PRS had an area under the receiver operating characteristic curve of 0.606. The lifetime risk of breast cancer for Shanghai Chinese women in the lowest and highest 1% of the PRS was 1.35% and 10.06%, respectively. CONCLUSION: Approximately one-half of GWAS-identified breast cancer risk variants can be directly replicated in East Asian women. Collectively, common genetic variants are important predictors for breast cancer risk. Using common genetic variants for breast cancer could help identify women at high risk of breast cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".