Genome-wide association studies and Mendelian randomization analyses provide insights into the causes of early-onset colorectal cancer
Bibliographic record
Abstract
BACKGROUND: The incidence of early-onset colorectal cancer (EOCRC; diagnosed <50 years of age) is rising globally; however, the causes underlying this trend are largely unknown. CRC has strong genetic and environmental determinants, yet common genetic variants and causal modifiable risk factors underlying EOCRC are unknown. We conducted the first EOCRC-specific genome-wide association study (GWAS) and Mendelian randomization (MR) analyses to explore germline genetic and causal modifiable risk factors associated with EOCRC. PATIENTS AND METHODS: We conducted a GWAS meta-analysis of 6176 EOCRC cases and 65 829 controls from the Genetics and Epidemiology of Colorectal Cancer Consortium (GECCO), the Colorectal Transdisciplinary Study (CORECT), the Colon Cancer Family Registry (CCFR), and the UK Biobank. We then used the EOCRC GWAS to investigate 28 modifiable risk factors using two-sample MR. RESULTS: We found two novel risk loci for EOCRC at 1p34.1 and 4p15.33, which were not previously associated with CRC risk. We identified a deleterious coding variant (rs36053993, G396D) at polyposis-associated DNA repair gene MUTYH (odds ratio 1.80, 95% confidence interval 1.47-2.22) but show that most of the common genetic susceptibility was from noncoding signals enriched in epigenetic markers present in gastrointestinal tract cells. We identified new EOCRC-susceptibility genes, and in addition to pathways such as transforming growth factor (TGF) β, suppressor of Mothers Against Decapentaplegic (SMAD), bone morphogenetic protein (BMP) and phosphatidylinositol kinase (PI3K) signaling, our study highlights a role for insulin signaling and immune/infection-related pathways in EOCRC. In our MR analyses, we found novel evidence of probable causal associations for higher levels of body size and metabolic factors-such as body fat percentage, waist circumference, waist-to-hip ratio, basal metabolic rate, and fasting insulin-higher alcohol drinking, and lower education attainment with increased EOCRC risk. CONCLUSIONS: Our novel findings indicate inherited susceptibility to EOCRC and suggest modifiable lifestyle and metabolic targets that could also be used to risk-stratify individuals for personalized screening strategies or other interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.016 | 0.039 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".