A region-based gene association study combined with a leave-one-out sensitivity analysis identifies SMG1 as a pancreatic cancer susceptibility gene
Bibliographic record
Abstract
Pancreatic adenocarcinoma (PC) is a lethal malignancy that is familial or associated with genetic syndromes in 10% of cases. Gene-based surveillance strategies for at-risk individuals may improve clinical outcomes. However, familial PC (FPC) is plagued by genetic heterogeneity and the genetic basis for the majority of FPC remains elusive, hampering the development of gene-based surveillance programs. The study was powered to identify genes with a cumulative pathogenic variant prevalence of at least 3%, which includes the most prevalent PC susceptibility gene, BRCA2. Since the majority of known PC susceptibility genes are involved in DNA repair, we focused on genes implicated in these pathways. We performed a region-based association study using the Mixed-Effects Score Test, followed by leave-one-out characterization of PC-associated gene regions and variants to identify the genes and variants driving risk associations. We evaluated 398 cases from two case series and 987 controls without a personal history of cancer. The first case series consisted of 109 patients with either FPC (n = 101) or PC at ≤50 years of age (n = 8). The second case series was composed of 289 unselected PC cases. We validated this discovery strategy by identifying known pathogenic BRCA2 variants, and also identified SMG1, encoding a serine/threonine protein kinase, to be significantly associated with PC following correction for multiple testing (p = 3.22x10-7). The SMG1 association was validated in a second independent series of 532 FPC cases and 753 controls (p<0.0062, OR = 1.88, 95%CI 1.17-3.03). We showed segregation of the c.4249A>G SMG1 variant in 3 affected relatives in a FPC kindred, and we found c.103G>A to be a recurrent SMG1 variant associating with PC in both the discovery and validation series. These results suggest that SMG1 is a novel PC susceptibility gene, and we identified specific SMG1 gene variants associated with PC risk.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".