Socioeconomic Status, Smoking, and Lung Cancer: Mediation and Bias Analysis in the SYNERGY Study
Bibliographic record
Abstract
BACKGROUND: Increased lung cancer risks for low socioeconomic status (SES) groups are only partially attributable to smoking habits. Little effort has been made to investigate the persistent risks related to low SES by quantification of potential biases. METHODS: Based on 12 case-control studies, including 18 centers of the international SYNERGY project (16,550 cases, 20,147 controls), we estimated controlled direct effects (CDE) of SES on lung cancer via multiple logistic regression, adjusted for age, study center, and smoking habits and stratified by sex. We conducted mediation analysis by inverse odds ratio weighting to estimate natural direct effects and natural indirect effects via smoking habits. We considered misclassification of smoking status, selection bias, and unmeasured mediator-outcome confounding by genetic risk, both separately and by multiple quantitative bias analyses, using bootstrap to create 95% simulation intervals (SI). RESULTS: Mediation analysis of lung cancer risks for SES estimated mean proportions of 43% in men and 33% in women attributable to smoking. Bias analyses decreased the direct effects of SES on lung cancer, with selection bias showing the strongest reduction in lung cancer risk in the multiple bias analysis. Lung cancer risks remained increased for lower SES groups, with higher risks in men (fourth vs. first [highest] SES quartile: CDE, 1.50 [SI, 1.32, 1.69]) than women (CDE: 1.20 [SI: 1.01, 1.45]). Natural direct effects were similar to CDE, particularly in men. CONCLUSIONS: Bias adjustment lowered direct lung cancer risk estimates of lower SES groups. However, risks for low SES remained elevated, likely attributable to occupational hazards or other environmental exposures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.080 | 0.129 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.008 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".