Bayesian modelling of lung cancer risk and bitumen fume exposure adjusted for unmeasured confounding by smoking
Bibliographic record
Abstract
OBJECTIVES: Residual confounding can be present in epidemiological studies because information on confounding factors was not collected. A Bayesian framework, which has the advantage over frequentist methods that the uncertainty in the association between the confounding factor and exposure and disease can be reflected in the credible intervals of the risk parameter, is proposed to assess the magnitude and direction of this bias. METHODS: To illustrate this method, bias from smoking as an unmeasured confounder in a cohort study of lung cancer risk in the European asphalt industry was assessed. A Poisson disease model was specified to assess lung cancer risk associated with career average, cumulative and lagged bitumen fume exposure. Prior distributions for the exposure strata, as well as for other covariates, were specified as uninformative normal distributions. The priors on smoking habits were specified as Dirichlet distributions based on smoking prevalence estimates available for a sub-cohort and assumptions about precision of these estimates. RESULTS: Median bias in this example was estimated at 13%, and suggested an attenuating effect on the original exposure-disease associations. Nonetheless, the results still implied an increased lung cancer risk, especially for average exposure. CONCLUSIONS: This Bayesian framework provides a method to assess the bias from an unmeasured confounding factor taking into account the uncertainty surrounding the estimate and from random sampling error. Specifically for this example, the bias arising from unmeasured smoking history in this asphalt workers' cohort is unlikely to explain the increased lung cancer risk associated with average bitumen fume exposure found in the original study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".