Accounting for Intra-urban Variability in Outdoor Air Concentrations: Estimating Exposures Using Monitoring Station Data and Land Use Regression Models
Bibliographic record
Abstract
P-640 Introduction: Several methods are available for estimating outdoor air pollution concentrations at specific residential locations as a means of estimating personal exposures in epidemiological studies or risk assessment. We compared several such methods used in the Border Air Quality Study. Methods: One approach uses data from regulatory ambient monitoring stations. Examples include assigning individuals the concentration of (1) the nearest monitor, (2) the average of all monitors within a specific distance, (3) the average of the 3 closest monitors. Inverse-distance-weighting and Kriging approaches are also possible with these data. A second approach uses spatially resolved air pollution models. Here, we focus on a land use regression model that relates measured concentrations to geographic data (e.g., traffic and road length) for uniform buffers surrounding monitoring sites. The model is then used to predict outdoor concentrations at specific locations. Outdoor concentration estimates were generated using each of the above methods and compared for 62,531 postal codes in the study area, and for each of 114,400 pregnant women in a birth outcome cohort study. Seven pollutants are considered: NO, NO2, ozone, CO, SO2, PM10, and PM2.5. Results: The various approaches considered are often poorly correlated with each other (R values range: 0.06–0.99; median: 0.63), both for individuals in the cohort and for all postal codes in the study area. Concentration maps highlight differences among the methods. Visual inspection of these maps, comparing attributes such as spatial autocorrelation and spatial heterogeneity, suggests that in some cases, methods that seem reasonable in theory (and have been implemented in previous epidemiological studies) yield results that seem unlikely to be the best estimate of the expected concentration. The ambient monitoring data are spatially sparse (monitor density, in km2 per monitor: 120 for PM2.5; 12–24 for the remaining six pollutants) relative to the land-use regression model (calculated on a 10-meter grid, for NO and NO2). Monitoring station data have several unique statistical properties (e.g., length scales for spatial heterogeneity vary among pollutants and locations, and over time); for the case studied, these properties suggest that the inverse-distance-weighted average concentration from the 3 nearest monitoring stations is preferable to the alternatives considered. Discussion: Epidemiological study results may vary, depending on which exposure method is selected. Researchers should be aware of potential biases in their studies due to the choice of exposure method.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".