Abstract 229: Genome-wide association study by colorectal carcinoma subtype
Bibliographic record
Abstract
Abstract Over 50 genetic variants have been associated with colorectal cancer (CRC) risk through genome-wide association studies (GWAS), yet these variants represent only a fraction of the total estimated heritability. CRC is a heterogenous disease with diverse tumor etiology. Assessing genetic risk in molecular subtypes may help to identify novel loci and characterize genetic risk among tumor subtypes. We used microsatellite instability (MSI), an established CRC classifier with etiological and therapeutic relevance, to define CRC subtypes for GWAS analyses. We conducted a case-case analysis to estimate odds ratios (OR) and 95% confidence intervals (CI) for association of genome-wide variants with microsatellite stable (MSS) versus unstable (MSI) carcinomas. We ran an inverse-variance weighted fixed-effects meta-analysis across GWAS in a discovery set of 4,163 population-based CRC cases with harmonized microsatellite instability (MSI) marker and imputed genotype data. For each analysis, we used log-additive logistic regression, adjusting for age, sex, and principal components to account for population substructure. We then followed up with replication of 102 SNPs that reached p-values less than 5x10-6 in 1,698 cases. A total of 845 (20.3%) cancer cases were microsatellite unstable in the discovery population and 174 (10.2%) were unstable in the replication population. No variants reached the genome-wide significance level of 5x10-8 in the discovery set. However, we identified two variants that reached a Bonferroni corrected p-value of 4.0x10-4 in the replication set. This included one variant in MLH1 (Replication: OR=1.74, 95% CI=1.53-1.98, p=1.63x10-5; Discovery+Replication: OR=1.45, 95% CI=1.37-1.54, p=9.76x10-11) and one variant in LOC105377645 (Replication: OR=1.70, 95% CI=1.49-1.94, p=5.13x10-5; Discovery+Replication: OR=1.45, 95% CI=1.37-1.54, p=9.76 x 10-11). The MLH1 gene is a DNA mismatch repair gene implicated in Lynch Syndrome, the hallmark of which is microsatellite instability. This is the first genome-wide scan to identify a common variant in MLH1 that is associated with CRC. This variant (minor allele frequency, MAF = 23% in this all European ancestry population) is located in the 5'-untranslated region of MLH1 and is thought to act as a long-range regulator of DCLK3, a potential tumor driver gene. The second variant, located in LOC105377645 with an MAF of 22%, is in an uncharacterized region of the genome and has not previously been implicated in cancer development. These findings suggest that accounting for molecular heterogeneity is important for discovery and characterization of genetic variants associated with CRC risk. We plan to run polytomous regression analyses, increase our sample size, and further investigate CRC subtypes by CIMP, BRAF mutation, KRAS mutation status. Citation Format: Tabitha A. Harrison, Yiwen Lu, Chenjie Zeng, Flora Qu, Kristin Anderson, Hermann Brenner, Daniel D. Buchanan, Peter T. Campbell, Andrew T. Chan, Jenny Chang-Claude, Graham G. Giles, Bethany Van Guelpen, Michael Hoffmeister, Mark A. Jenkins, Noralane M. Lindor, Roger L. Milne, Polly A. Newcomb, Reiko Nishihara, Michael O. Woods, Shuji Ogino, John D. Potter, Martha L. Slattery, Wei Sun, Stephen N. Thibodeau, Li Hsu, Ulrike Peters. Genome-wide association study by colorectal carcinoma subtype [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 229.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.005 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".