Multi-ancestry GWAS meta-analyses of lung cancer reveal susceptibility loci and elucidate smoking-independent genetic risk
Bibliographic record
Abstract
Lung cancer remains the leading cause of cancer mortality, despite declining smoking rates. Previous lung cancer GWAS have identified numerous loci, but separating the genetic risks of lung cancer and smoking behavioral susceptibility remains challenging. Here, we perform multi-ancestry GWAS meta-analyses of lung cancer using the Million Veteran Program cohort (approximately 95% male cases) and a previous study of European-ancestry individuals, jointly comprising 42,102 cases and 181,270 controls, followed by replication in an independent cohort of 19,404 cases and 17,378 controls. We then carry out conditional meta-analyses on cigarettes per day and identify two novel, replicated loci, including the 19p13.11 pleiotropic cancer locus in squamous cell lung carcinoma. Overall, we report twelve novel risk loci for overall lung cancer, lung adenocarcinoma, and squamous cell lung carcinoma, nine of which are externally replicated. Finally, we perform PheWAS on polygenic risk scores for lung cancer, with and without conditioning on smoking. The unconditioned lung cancer polygenic risk score is associated with smoking status in controls, illustrating a reduced predictive utility in non-smokers. Additionally, our polygenic risk score demonstrates smoking-independent pleiotropy of lung cancer risk across neoplasms and metabolic traits.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".