Socioeconomic Status, Smoking, and Lung Cancer: Mediation and Bias Analysis in the SYNERGY Study
Bibliographic record
Abstract
BACKGROUND: Increased lung cancer risks for low socioeconomic status (SES) groups are only partially attributable to smoking habits. Little effort has been made to investigate the persistent risks related to low SES by quantification of potential biases. METHODS: Based on 12 case-control studies, including 18 centers of the international SYNERGY project (16,550 cases, 20,147 controls), we estimated controlled direct effects (CDE) of SES on lung cancer via multiple logistic regression, adjusted for age, study center, and smoking habits and stratified by sex. We conducted mediation analysis by inverse odds ratio weighting to estimate natural direct effects and natural indirect effects via smoking habits. We considered misclassification of smoking status, selection bias, and unmeasured mediator-outcome confounding by genetic risk, both separately and by multiple quantitative bias analyses, using bootstrap to create 95% simulation intervals (SI). RESULTS: Mediation analysis of lung cancer risks for SES estimated mean proportions of 43% in men and 33% in women attributable to smoking. Bias analyses decreased the direct effects of SES on lung cancer, with selection bias showing the strongest reduction in lung cancer risk in the multiple bias analysis. Lung cancer risks remained increased for lower SES groups, with higher risks in men (fourth vs. first [highest] SES quartile: CDE, 1.50 [SI, 1.32, 1.69]) than women (CDE: 1.20 [SI: 1.01, 1.45]). Natural direct effects were similar to CDE, particularly in men. CONCLUSIONS: Bias adjustment lowered direct lung cancer risk estimates of lower SES groups. However, risks for low SES remained elevated, likely attributable to occupational hazards or other environmental exposures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".