Interaction between Continuous Pack-Years Smoked and Polygenic Risk Score on Lung Cancer Risk: Prospective Results from the Framingham Heart Study
Bibliographic record
Abstract
BACKGROUND: Lung cancer risk attributable to smoking is dose dependent, yet few studies examining a polygenic risk score (PRS) by smoking interaction have included comprehensive lifetime pack-years smoked. METHODS: We analyzed data from participants of European ancestry in the Framingham Heart Study Original (n = 454) and Offspring (n = 2,470) cohorts enrolled in 1954 and 1971, respectively, and followed through 2018. We built a PRS for lung cancer using participant genotyping data and genome-wide association study summary statistics from a recent study in the OncoArray Consortium. We used Cox proportional hazards regression models to assess risk and the interaction between pack-years smoked and genetic risk for lung cancer adjusting for European ancestry, age, sex, and education. RESULTS: We observed a significant submultiplicative interaction between pack-years and PRS on lung cancer risk (P = 0.09). Thus, the relative risk associated with each additional 10 pack-years smoked decreased with increasing genetic risk (HR = 1.56 at one SD below mean PRS, HR = 1.48 at mean PRS, and HR = 1.40 at one SD above mean PRS). Similarly, lung cancer risk per SD increase in the PRS was highest among those who had never smoked (HR = 1.55) and decreased with heavier smoking (HR = 1.32 at 30 pack-years). CONCLUSIONS: These results suggest the presence of a submultiplicative interaction between pack-years and genetics on lung cancer risk, consistent with recent findings. Both smoking and genetics were significantly associated with lung cancer risk. IMPACT: These results underscore the contributions of genetics and smoking on lung cancer risk and highlight the negative impact of continued smoking regardless of genetic risk.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".