Interactions between folate intake and genetic predictors of gene expression levels associated with colorectal cancer risk
Bibliographic record
Abstract
Observational studies have shown higher folate consumption to be associated with lower risk of colorectal cancer (CRC). Understanding whether and how genetic risk factors interact with folate could further elucidate the underlying mechanism. Aggregating functionally relevant genetic variants in set-based variant testing has higher power to detect gene-environment (G × E) interactions and may provide information on the underlying biological pathway. We investigated interactions between folate consumption and predicted gene expression on colorectal cancer risk across the genome. We used variant weights from the PrediXcan models of colon tissue-specific gene expression as a priori variant information for a set-based G × E approach. We harmonized total folate intake (mcg/day) based on dietary intake and supplemental use across cohort and case-control studies and calculated sex and study specific quantiles. Analyses were performed using a mixed effects score tests for interactions between folate and genetically predicted expression of 4839 genes with available genetically predicted expression. We pooled results across 23 studies for a total of 13,498 cases with colorectal tumors and 13,918 controls of European ancestry. We used a false discovery rate of 0.2 to identify genes with suggestive evidence of an interaction. We found suggestive evidence of interaction with folate intake on CRC risk for genes including glutathione S-Transferase Alpha 1 (GSTA1; p = 4.3E-4), Tonsuko Like, DNA Repair Protein (TONSL; p = 4.3E-4), and Aspartylglucosaminidase (AGA: p = 4.5E-4). We identified three genes involved in preventing or repairing DNA damage that may interact with folate consumption to alter CRC risk. Glutathione is an antioxidant, preventing cellular damage and is a downstream metabolite of homocysteine and metabolized by GSTA1. TONSL is part of a complex that functions in the recovery of double strand breaks and AGA plays a role in lysosomal breakdown of glycoprotein.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".