Investigation of gene‐environment interactions between 47 newly identified breast cancer susceptibility loci and environmental risk factors
Bibliographic record
Abstract
A large genotyping project within the Breast Cancer Association Consortium (BCAC) recently identified 41 associations between single nucleotide polymorphisms (SNPs) and overall breast cancer (BC) risk. We investigated whether the effects of these 41 SNPs, as well as six SNPs associated with estrogen receptor (ER) negative BC risk are modified by 13 environmental risk factors for BC. Data from 22 studies participating in BCAC were pooled, comprising up to 26,633 cases and 30,119 controls. Interactions between SNPs and environmental factors were evaluated using an empirical Bayes-type shrinkage estimator. Six SNPs showed interactions with associated p-values (pint ) <1.1 × 10(-3) . None of the observed interactions was significant after accounting for multiple testing. The Bayesian False Discovery Probability was used to rank the findings, which indicated three interactions as being noteworthy at 1% prior probability of interaction. SNP rs6828523 was associated with increased ER-negative BC risk in women ≥170 cm (OR = 1.22, p = 0.017), but inversely associated with ER-negative BC risk in women <160 cm (OR = 0.83, p = 0.039, pint = 1.9 × 10(-4) ). The inverse association between rs4808801 and overall BC risk was stronger for women who had had four or more pregnancies (OR = 0.85, p = 2.0 × 10(-4) ), and absent in women who had had just one (OR = 0.96, p = 0.19, pint = 6.1 × 10(-4) ). SNP rs11242675 was inversely associated with overall BC risk in never/former smokers (OR = 0.93, p = 2.8 × 10(-5) ), but no association was observed in current smokers (OR = 1.07, p = 0.14, pint = 3.4 × 10(-4) ). In conclusion, recently identified BC susceptibility loci are not strongly modified by established risk factors and the observed potential interactions require confirmation in independent studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".