Beyond GWAS of Colorectal Cancer: Evidence of Interaction with Alcohol Consumption and Putative Causal Variant for the 10q24.2 Region
Bibliographic record
Abstract
BACKGROUND: Currently known associations between common genetic variants and colorectal cancer explain less than half of its heritability of 25%. As alcohol consumption has a J-shape association with colorectal cancer risk, nondrinking and heavy drinking are both risk factors for colorectal cancer. METHODS: Individual-level data was pooled from the Colon Cancer Family Registry, Colorectal Transdisciplinary Study, and Genetics and Epidemiology of Colorectal Cancer Consortium to compare nondrinkers (≤1 g/day) and heavy drinkers (>28 g/day) with light-to-moderate drinkers (1-28 g/day) in GxE analyses. To improve power, we implemented joint 2df and 3df tests and a novel two-step method that modifies the weighted hypothesis testing framework. We prioritized putative causal variants by predicting allelic effects using support vector machine models. RESULTS: For nondrinking as compared with light-to-moderate drinking, the hybrid two-step approach identified 13 significant SNPs with pairwise r2 > 0.9 in the 10q24.2/COX15 region. When stratified by alcohol intake, the A allele of lead SNP rs2300985 has a dose-response increase in risk of colorectal cancer as compared with the G allele in light-to-moderate drinkers [OR for GA genotype = 1.11; 95% confidence interval (CI), 1.06-1.17; OR for AA genotype = 1.22; 95% CI, 1.14-1.31], but not in nondrinkers or heavy drinkers. Among the correlated candidate SNPs in the 10q24.2/COX15 region, rs1318920 was predicted to disrupt an HNF4 transcription factor binding motif. CONCLUSIONS: Our study suggests that the association with colorectal cancer in 10q24.2/COX15 observed in genome-wide association study is strongest in nondrinkers. We also identified rs1318920 as the putative causal regulatory variant for the region. IMPACT: The study identifies multifaceted evidence of a possible functional effect for rs1318920.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".