Genome-wide association study provides evidence for a breast cancer risk locus at 6q22.33
Bibliographic record
Abstract
We performed a three-phase genome-wide association study (GWAS) using cases and controls from a genetically isolated population, Ashkenazi Jews (AJ), to identify loci associated with breast cancer risk. In the first phase, we compared allele frequencies of 150,080 SNPs in 249 high-risk, BRCA1/2 mutation-negative AJ familial cases and 299 cancer-free AJ controls using chi(2) and the Cochran-Armitage trend tests. In the second phase, we genotyped 343 SNPs from 123 regions most significantly associated from stage 1, including 4 SNPs from the FGFR2 region, in 950 consecutive AJ breast cancer cases and 979 age-matched AJ controls. We replicated major associations in a third independent set of 243 AJ cases and 187 controls. We obtained a significant allele P value of association with AJ breast cancer in the FGFR2 region (P = 1.5 x 10(-5), odds ratio (OR) 1.26, 95% confidence interval (CI) 1.13-1.40 at rs1078806 for all phases combined). In addition, we found a risk locus in a region of chromosome 6q22.33 (P = 2.9 x 10(-8), OR 1.41, 95% CI 1.25-1.59 at rs2180341). Using several SNPs at each implicated locus, we were able to verify associations and impute haplotypes. The major haplotype at the 6q22.33 locus conferred protection from disease, whereas the minor haplotype conferred risk. Candidate genes in the 6q22.33 region include ECHDC1, which encodes a protein involved in mitochondrial fatty acid oxidation, and also RNF146, which encodes a ubiquitin protein ligase, both known pathways in breast cancer pathogenesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".