Comparing land use regression and dispersion modelling to assess residential exposure to ambient air pollution for epidemiological studies
Bibliographic record
Abstract
BACKGROUND: Land-use regression (LUR) and dispersion models (DM) are commonly used for estimating individual air pollution exposure in population studies. Few comparisons have however been made of the performance of these methods. OBJECTIVES: Within the European Study of Cohorts for Air Pollution Effects (ESCAPE) we explored the differences between LUR and DM estimates for NO2, PM10 and PM2.5. METHODS: The ESCAPE study developed LUR models for outdoor air pollution levels based on a harmonised monitoring campaign. In thirteen ESCAPE study areas we further applied dispersion models. We compared LUR and DM estimates at the residential addresses of participants in 13 cohorts for NO2; 7 for PM10 and 4 for PM2.5. Additionally, we compared the DM estimates with measured concentrations at the 20-40 ESCAPE monitoring sites in each area. RESULTS: The median Pearson R (range) correlation coefficients between LUR and DM estimates for the annual average concentrations of NO2, PM10 and PM2.5 were 0.75 (0.19-0.89), 0.39 (0.23-0.66) and 0.29 (0.22-0.81) for 112,971 (13 study areas), 69,591 (7) and 28,519 (4) addresses respectively. The median Pearson R correlation coefficients (range) between DM estimates and ESCAPE measurements were of 0.74 (0.09-0.86) for NO2; 0.58 (0.36-0.88) for PM10 and 0.58 (0.39-0.66) for PM2.5. CONCLUSIONS: LUR and dispersion model estimates correlated on average well for NO2 but only moderately for PM10 and PM2.5, with large variability across areas. DM predicted a moderate to large proportion of the measured variation for NO2 but less for PM10 and PM2.5.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.059 | 0.127 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.005 |
| Bibliometrics | 0.004 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".