Molecular Signature of Smoking in Human Lung Tissues
Bibliographic record
Abstract
Cigarette smoking is the leading risk factor for lung cancer. To identify genes deregulated by smoking and to distinguish gene expression changes that are reversible and persistent following smoking cessation, we carried out genome-wide gene expression profiling on nontumor lung tissue from 853 patients with lung cancer. Gene expression levels were compared between never and current smokers, and time-dependent changes in gene expression were studied in former smokers. A total of 3,223 transcripts were differentially expressed between smoking groups in the discovery set (n = 344, P < 1.29 × 10(-6)). A substantial number of smoking-induced genes also were validated in two replication sets (n = 285 and 224), and a gene expression signature of 599 transcripts consistently segregated never from current smokers across all three sets. The expression of the majority of these genes reverted to never-smoker levels following smoking cessation, although the time course of normalization differed widely among transcripts. Moreover, some genes showed very slow or no reversibility in expression, including SERPIND1, which was found to be the most consistent gene permanently altered by smoking in the three sets. Our findings therefore indicate that smoking deregulates many genes, many of which reverse to normal following smoking cessation. However, a subset of genes remains altered even decades following smoking cessation and may account, at least in part, for the residual risk of lung cancer among former smokers. Cancer Res; 72(15); 3753-63. ©2012 AACR.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".