Genetic variant in DNA repair gene<i>GTF2H4</i>is associated with lung cancer risk: a large-scale analysis of six published GWAS datasets in the TRICL consortium
Bibliographic record
Abstract
DNA repair pathways maintain genomic integrity and stability, and dysfunction of DNA repair leads to cancer. We hypothesize that functional genetic variants in DNA repair genes are associated with risk of lung cancer. We performed a large-scale meta-analysis of 123,371 single nucleotide polymorphisms (SNPs) in 169 DNA repair genes obtained from six previously published genome-wide association studies (GWASs) of 12160 lung cancer cases and 16838 controls. We calculated odds ratios (ORs) with 95% confidence intervals (CIs) using the logistic regression model and used the false discovery rate (FDR) method for correction of multiple testing. As a result, 14 SNPs had a significant odds ratio (OR) for lung cancer risk with P FDR < 0.05, of which rs3115672 in MSH5 (OR = 1.20, 95% CI = 1.14-1.27) and rs114596632 in GTF2H4 (OR = 1.19, 95% CI = 1.12-1.25) at 6q21.33 were the most statistically significant (P combined = 3.99×10(-11) and P combined = 5.40×10(-10), respectively). The MSH5 rs3115672, but not GTF2H4 rs114596632, was strongly correlated with MSH5 rs3131379 in that region (r (2) = 1.000 and r (2) = 0.539, respectively) as reported in a previous GWAS. Importantly, however, the GTF2H4 rs114596632 T, but not MSH5 rs3115672 T, allele was significantly associated with both decreased DNA repair capacity phenotype and decreased mRNA expression levels. These provided evidence that functional genetic variants of GTF2H4 confer susceptibility to lung cancer.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.005 |
| Bibliometrics | 0.002 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".