Genome-wide association study for lung cancer in 6531 African Americans reveals new susceptibility loci
Bibliographic record
Abstract
Despite lung cancer affecting all races and ethnicities, disparities are observed in incidence and mortality rates among different ethnic groups in the United States. Non-Hispanic African Americans had a high incidence rate of lung cancer at 55.8 per 100 000 people, as well as the highest death rate at 37.2 per 100 000 people from 2016 to 2020. While previous genome-wide association studies (GWAS) have identified over 45 susceptibility risk loci that influence lung cancer development, few GWAS have investigated the etiology of lung cancer in African Americans. To address this gap in knowledge, we conducted GWAS of lung cancer focused on studying African Americans, comprising 2267 lung cancer cases and 4264 controls. We identified three loci associated with lung cancer, one with lung adenocarcinoma, and four with lung squamous cell carcinoma in this population at the genomic-wide significance level. Among them, three novel loci were identified near VWF at 12p13.31 for overall lung cancer and GACAT3 at 2p24.3 and LMAN1L at 15q24.1 for lung squamous cell carcinoma. In addition, we confirmed previously reported risk loci with known or new lead variants near CHRNA5 at 15q25.1 and CYP2A6 at 19q13.2 associated with lung cancer and TRIP13 at 5p15.33 and ERC1 at 12p13.33 associated with lung squamous cell carcinoma. Further multi-step functional analyses shed light on biological mechanisms underlying these associations of lung cancer in this population. Our study highlights the importance of ancestry-specific studies for the potential alleviation of lung cancer burden in African Americans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".