Benzene exposure assessment uncertainty and its use in sensitivity analyses in a pooled study of leukaemia
Bibliographic record
Abstract
Objectives We examined the effect of exposure misclassification on measures of risk in a study of leukaemia and exposure to benzene. Methods The exposure to benzene was estimated for 370 leukaemia cases and 1587 controls in a nested case-control study. The study drew cases and controls from three petroleum industry cohorts in the UK, Canada and Australia. The exposures were expressed as intensity of benzene exposure (ppm) for each job and the cumulative exposure (ppm-years) was estimated by multiplying the years at each job and summing over the career. The exposure estimate for each job was allocated a confidence 25ore of Low, Medium or High. This was used to group the subjects into those for whom all jobs with a High confidence exposure estimates, and those with a Moderate confidence exposure estimates (all jobs allocated High or Medium ie, where no jobs had a Low score). The ORs for leukaemia subtypes associated with cumulative exposure or intensity of exposure were estimated. Sensitivity analyses included only those subjects with Moderate or High confidence scores. Results For more than one leukaemia subtypes the ORs were increased when subjects with low exposure certainty were removed and further increased when only those subjects with high certainty were included. Conclusions This suggests that exposure misclassification has reduced the observed ORs in this study and highlights the need to design studies so that this can be evaluated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.421 | 0.627 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.004 | 0.020 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".