The spatial association between community air pollution and mortality: a new method of analyzing correlated geographic cohort data.
Bibliographic record
Abstract
We present a new statistical model for linking spatial variation in ambient air pollution to mortality. The model incorporates risk factors measured at the individual level, such as smoking, and at the spatial level, such as air pollution. We demonstrate that the spatial autocorrelation in community mortality rates, an indication of not fully characterizing potentially confounding risk factors to the air pollution-mortality association, can be accounted for through the inclusion of location in the model assessing the effects of air pollution on mortality. Our methods are illustrated with an analysis of the American Cancer Society cohort to determine whether all cause mortality is associated with concentrations of sulfate particles. The relative risk associated with a 4.2 microg/m(3) interquartile range of sulfate distribution for all causes of death was 1.051 (95% confidence interval 1.036-1.066) based on the Cox proportional hazards survival model, assuming subjects were statistically independent. Inclusion of community-based random effects yielded a relative risk of 1.055 (1.033, 1.077), which represented a doubling in the residual variance compared to that estimated by the Cox model. Residuals from the random-effects model displayed strong evidence of spatial autocorrelation (p = 0.0052). Further inclusion of a location surface reduced the sulfate relative risk and the evidence for autocorrelation as the complexity of the location surface increased, with a range in relative risks of 1.055-1.035. We conclude that these data display both extravariation and spatial autocorrelation, characteristics not captured by the Cox survival model. Failure to account for extravariation and spatial autocorrelation can lead to an understatement of the uncertainty of the air pollution association with mortality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".