Novel insights into genetic susceptibility for colorectal cancer from transcriptome-wide association and functional investigation
Bibliographic record
Abstract
BACKGROUND: Transcriptome-wide association studies have been successful in identifying candidate susceptibility genes for colorectal cancer (CRC). To strengthen susceptibility gene discovery, we conducted a large transcriptome-wide association study and an alternative splicing transcriptome-wide association study in CRC using improved genetic prediction models and performed in-depth functional investigations. METHODS: We analyzed RNA-sequencing data from normal colon tissues and genotype data from 423 European descendants to build genetic prediction models of gene expression and alternative splicing and evaluated model performance using independent RNA-sequencing data from normal colon tissues of the Genotype-Tissue Expression Project. We applied the verified models to genome-wide association studies (GWAS) summary statistics among 58 131 CRC cases and 67 347 controls of European ancestry to evaluate associations of genetically predicted gene expression and alternative splicing with CRC risk. We performed in vitro functional assays for 3 selected genes in multiple CRC cell lines. RESULTS: We identified 57 putative CRC susceptibility genes, which included the 48 genes from transcriptome-wide association studies and 15 genes from splicing transcriptome-wide association studies, at a Bonferroni-corrected P value less than .05. Of these, 16 genes were not previously implicated in CRC susceptibility, including a gene PDE7B (6q23.3) at locus previously not reported by CRC GWAS. Gene knockdown experiments confirmed the oncogenic roles for 2 unreported genes, TRPS1 and METRNL, and a recently reported gene, C14orf166. CONCLUSION: This study discovered new putative susceptibility genes of CRC and provided novel insights into the biological mechanisms underlying CRC development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".