Is the impact of cigarette smoking on lung cancer risk different between males and females?
Bibliographic record
Abstract
Lung cancer is the leading cause of cancer-related mortality throughout the world and cigarette smoking is the most important risk factor. A topic of considerable interest from both etiologic and public health perspectives is whether women have different susceptibility to smoking-induced lung cancer than men. Published epidemiologic studies have produced discrepant evidence. The discrepancies may partly be due to methodological considerations. In this thesis, I first investigated some relevant methodological issues, and then estimated and tested the sex*smoking interaction using a large case-control study of lung cancer. Investigation of potentially differential sex susceptibility to smoking carcinogenesis can be cast as a problem of assessing sex*smoking interaction. One issue that arises when interpreting the results from past studies is the estimated sex*smoking interaction effects might be biased due to unobserved confounding. Therefore, in the 1st manuscript I investigated the conditions under which an interaction effect estimate will be confounded. I identified two different situations where the failure to adjust for the effect of a risk factor U results in a biased estimate of the interaction, assessed on multiplicative scales, between exposures E1 and E2 on a binary outcome Y: (1) U is associated with E1 and has an interaction with E2 for Y; (2) the association between U and E1 varies depending on the value of E2. Investigation of potential sex*smoking interaction should also consider potentially differential effects of continuous measures of smoking history. Although smoking intensity and cumulative exposure have been shown to have non-linear effects on the logit of lung cancer risk, previous studies that assessed their interactions with sex have a priori assumed their effects are linear. Thus, in the 2nd manuscript, I used simulations to assess the impact of mis-modeling non-linear effect of a continuous exposure on testing its multiplicative interaction with a binary covariate, an issue that has not yet been systematically investigated in statistical and epidemiological literature. The results indicate that mis-modeling the non-linear effect with the conventional linear function will result in an inflated type I error rate for the interaction test, only if the distribution of the continuous variable varies across the strata of the covariate in the interaction term. In the 3rd manuscript, I assessed whether cigarette smoking had a different impact on lung cancer risk between males and females using data from a large population-based case-control study conducted in 1996-2002 in Montreal. Multivariable logistic regression was used to assess the multiplicative interaction between sex and different smoking indices. To overcome limitations of some previous epidemiologic studies we adjusted for important confounders, and modeled more accurately the non-linear effects of different components of smoking history. The results indicate an interaction between sex and the binary indicator of ever smoking, with females showing significantly higher impact of cigarette smoking on lung cancer risk. The effects of smoking intensity and cumulative smoking exposure were stronger for female smokers than male smokers, although the respective interactions were non-significant. However, the impact of Comprehensive Smoking Index, a single aggregated measure of smoking exposure, was significantly stronger among the female smokers. Overall, the results of my thesis add new evidence to support the hypothesis that females are more susceptible to cigarette-induced lung cancer than males. Given the greater uptake of smoking in recent decades by women than by men, it implies that more vigorous efforts should be directed at eliminating smoking among women. Methodological contributions of my thesis will help enhance the validity and accuracy of epidemiological studies of many other interactions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".