Abstract 7397: A multi-ancestry polygenic risk score improves lung cancer risk stratification across diverse populations
Bibliographic record
Abstract
Abstract Background: Currently implemented lung cancer risk prediction models consider a limited set of risk factors and insufficiently stratify patients for screening, leaving a large proportion of lung cancer patients, particularly those from non-European ancestries, ineligible for low-dose CT scans. Polygenic risk scores (PRS) have demonstrated potential for improving lung cancer risk prediction, but are predominantly derived from participants of European-ancestry, limiting their applicability to other populations. This study aims to develop a novel multi-ancestry PRS for lung cancer and evaluate its ability to enhance equitable risk stratification. Methods: We constructed multi-ancestry PRS using the largest available lung cancer genome-wide associate studies (GWAS), which include 33, 023 European cases and 339, 471 controls, 11, 506 East Asian cases and 179, 654 controls, and 2, 379 African American cases and 6, 908 controls. We trained four genome-wide multi-ancestry PRS models (PRS-CSx, JointPRS, CT-SLEB, and PROSPER) using a reference panel of 1, 287, 077 SNPs, and compared them to a previously published 128-SNP PRS based on known lung cancer risk loci. PRS models were validated in 1, 047 lung cancer cases and 214, 403 controls from the All of Us Research Program biobank and their performance was evaluated in a pooled multi-ancestry analysis as well as stratified by genetic ancestry. Results: PRS-CSx demonstrated the highest performance among all tested methods, with an adjusted-AUC (conditional on sex, age, and the top 16 principal components) of 0.60 (95% CI: 0.58-0.62) and an OR of 1.43 (95% CI: 1.35-1.52) in the pooled analysis. Comparable performance was observed in the European ancestry subgroup (OR: 1.48, 95% CI: 1.37-1.59, 697 cases, 112, 145 controls). The PRS-CSx PRS also performed well in the African American population (OR 1.36, 95% CI: 1.17-1.58, 178 cases, 47, 701 controls), and in the Admixed American/Latino population (OR: 1.34, 95% CI: 1.03-1.74, 55 cases, 32, 936 controls). Individuals in the top decile of the PRS distribution had a 1.96-fold increased risk of lung cancer (95% CI: 1.60-2.40) compared to the average group in the 40-60th percentile. Furthermore, on average, individuals in the top decile reached a 1.5% 5-year absolute risk of lung cancer 8 years earlier than those with average PRS values. Conclusions: Multi-ancestry PRS outperform a single-ancestry PRS model. Integrating a multi-ancestry PRS into a lung cancer risk prediction model can improve risk stratification, addressing disparities in screening eligibility across diverse populations. This approach has the potential to improve early detection and reduce lung cancer mortality equitably across ancestry groups. Citation Format: Nina Adler, Tony Chen, Jinyoung Byun, Christopher Amos, Qing Lan, David C. Christiani, Mattias Johansson, James McKay, Maria T. Landi, Geoffrey Liu, Loic Le Marchand, The International Lung Cancer Consortium, Esteban J. Parra, Linda Kachuri, Haoyu Zhang, Rayjean J. Hung. A multi-ancestry polygenic risk score improves lung cancer risk stratification across diverse populations [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7397.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".