Spatiotemporal air pollution exposure assessment for a Canadian population-based lung cancer case-control study
Bibliographic record
Abstract
BACKGROUND: Few epidemiological studies of air pollution have used residential histories to develop long-term retrospective exposure estimates for multiple ambient air pollutants and vehicle and industrial emissions. We present such an exposure assessment for a Canadian population-based lung cancer case-control study of 8353 individuals using self-reported residential histories from 1975 to 1994. We also examine the implications of disregarding and/or improperly accounting for residential mobility in long-term exposure assessments. METHODS: National spatial surfaces of ambient air pollution were compiled from recent satellite-based estimates (for PM2.5 and NO2) and a chemical transport model (for O3). The surfaces were adjusted with historical annual air pollution monitoring data, using either spatiotemporal interpolation or linear regression. Model evaluation was conducted using an independent ten percent subset of monitoring data per year. Proximity to major roads, incorporating a temporal weighting factor based on Canadian mobile-source emission estimates, was used to estimate exposure to vehicle emissions. A comprehensive inventory of geocoded industries was used to estimate proximity to major and minor industrial emissions. RESULTS: Calibration of the national PM2.5 surface using annual spatiotemporal interpolation predicted historical PM2.5 measurement data best (R2 = 0.51), while linear regression incorporating the national surfaces, a time-trend and population density best predicted historical concentrations of NO2 (R2 = 0.38) and O3 (R2 = 0.56). Applying the models to study participants residential histories between 1975 and 1994 resulted in mean PM2.5, NO2 and O3 exposures of 11.3 μg/m3 (SD = 2.6), 17.7 ppb (4.1), and 26.4 ppb (3.4) respectively. On average, individuals lived within 300 m of a highway for 2.9 years (15% of exposure-years) and within 3 km of a major industrial emitter for 6.4 years (32% of exposure-years). Approximately 50% of individuals were classified into a different PM2.5, NO2 and O3 exposure quintile when using study entry postal codes and spatial pollution surfaces, in comparison to exposures derived from residential histories and spatiotemporal air pollution models. Recall bias was also present for self-reported residential histories prior to 1975, with cases recalling older residences more often than controls. CONCLUSIONS: We demonstrate a flexible exposure assessment approach for estimating historical air pollution concentrations over large geographical areas and time-periods. In addition, we highlight the importance of including residential histories in long-term exposure assessments. For submission to: Environmental Health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".