Genetic architectures of proximal and distal colorectal cancer are partly distinct
Bibliographic record
Abstract
ABSTRACT Objective An understanding of the etiologic heterogeneity of colorectal cancer (CRC) is critical for improving precision prevention, including individualized screening recommendations and the discovery of novel drug targets and repurposable drug candidates for chemoprevention. Known differences in molecular characteristics and environmental risk factors among tumors arising in different locations of the colorectum suggest partly distinct mechanisms of carcinogenesis. The extent to which the contribution of inherited genetic risk factors for sporadic CRC differs by anatomical subsite of the primary tumor has not been examined. Design To identify new anatomical subsite-specific risk loci, we performed genome-wide association study (GWAS) meta-analyses including data of 48,214 CRC cases and 64,159 controls of European ancestry. We characterized effect heterogeneity at CRC risk loci using multinomial modeling. Results We identified 13 loci that reached genome-wide significance (P <5×10 −8 ) and that were not reported by previous GWAS for overall CRC risk. Multiple lines of evidence support candidate genes at several of these loci. We detected substantial heterogeneity between anatomical subsites. Just over half (61) of 109 known and new risk variants showed no evidence for heterogeneity. In contrast, 22 variants showed association with distal CRC (including rectal cancer), but no evidence for association or an attenuated association with proximal CRC. For two loci, there was strong evidence for effects confined to proximal colon cancer. Conclusion Genetic architectures of proximal and distal CRC are partly distinct. Studies of risk factors and mechanisms of carcinogenesis, and precision prevention strategies should take into consideration the anatomical subsite of the tumor. Significance of this study What is already known about this subject? Heterogeneity among colorectal cancer (CRC) tumors originating at different locations of the colorectum has been revealed in somatic genomes, epigenomes, and transcriptomes, and in some established environmental risk factors for CRC. Genome-wide association studies (GWAS) have identified over 100 genetic variants for overall CRC risk; however, a comprehensive analysis of the extent to which genetic risk factors differ by the anatomical sublocation of the primary tumor is lacking. What are the new findings? In this large consortium-based study, we analyzed clinical and genome-wide genotype data of 112,373 CRC cases and controls of European ancestry to comprehensively examine whether CRC case subgroups defined by anatomical sublocation have distinct germline genetic etiologies. We discovered 13 new loci at genome-wide significance ( P <5×10 −8 ) that were specific to certain anatomical sublocations and that were not reported by previous GWAS for overall CRC risk; multiple lines of evidence support strong candidate target genes at several of these loci, including PTGER3, LCT, MLH1, CDX1, KLF14, PYGL, BCL11B , and BMP7 . Systematic heterogeneity analysis of genetic risk variants for CRC identified thus far, revealed that the genetic architectures of proximal and distal CRC are partly distinct. Taken together, our results further support the idea that tumors arising in different anatomical sublocations of the colorectum may have distinct etiologies. How might it impact on clinical practice in the foreseeable future? Our results provide an informative resource for understanding the differential role that genes and pathways may play in the mechanisms of proximal and distal CRC carcinogenesis. The new insights into the etiologies of proximal and distal CRC may inform the development of new precision prevention strategies, including individualized screening recommendations and the discovery of novel drug targets and repurposable drug candidates for chemoprevention. Our findings suggest that future studies of etiological risk factors for CRC and molecular mechanisms of carcinogenesis should take into consideration the anatomical sublocation of the colorectal tumor.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".