Fine‐mapping of a novel premenopausal breast cancer susceptibility locus at Chr4q31.22 in Caucasian women and validation in African and Chinese women
Bibliographic record
Abstract
We previously identified a novel breast cancer susceptibility variant on chromosome 4q31.22 locus (rs1429142) conferring risk among women of European ancestry. Here, we report replication of findings, validation of the variant in diverse populations and fine‐mapping of the associated locus in Caucasian population. The SNP rs1429142 (C/T, minor allele frequency 18%) showed association for the overall breast cancer risk in Stages 1–4 (n = 4,331 cases/4271 controls; p = 4.35 × 10−8; odds ratio, ORC‐allele,1.25), and an elevated risk among premenopausal women (n = 1,503 cases/4271 controls; p = 5.81 × 10−10; ORC‐allele 1.40) in European populations. SNP rs1429142 was associated with premenopausal breast cancer risk in women of African (T/C; p‐value 1.45 × 10−02; ORC‐allele 1.2) but not from Chinese ancestry. Fine‐mapping of the locus revealed several potential causal variants which are present within a single association signal, revealed from the conditional regression analysis. Functional annotation of the potential causal variants revealed three putative SNPs rs1366691, rs1429139 and rs7667633 with active enhancer functions inferred based on histone marks, DNase hypersensitive sites in breast cell line data. These putative variants were bound by transcription factors (C‐FOS, STAT1/3 and POL2/3) with known roles in inflammatory pathways. Furthermore, Hi‐C data revealed several short‐range interactions in the fine‐mapped locus harboring the putative variants. The fine mapped locus was predicted to be within a single topologically associated domain, potentially facilitating enhancer–promoter interactions possibly leading to the regulation of nearby genes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".