Methods, availability, and applications of PM<sub>2.5</sub> exposure estimates derived from ground measurements, satellite, and atmospheric models
Bibliographic record
Abstract
Fine particulate matter (PM2.5) is a well-established risk factor for public health. To support both health risk assessment and epidemiological studies, data are needed on spatial and temporal patterns of PM2.5 exposures. This review article surveys publicly available exposure datasets for surface PM2.5 mass concentrations over the contiguous U.S., summarizes their applications and limitations, and provides suggestions on future research needs. The complex landscape of satellite instruments, model capabilities, monitor networks, and data synthesis methods offers opportunities for research development, but would benefit from guidance for new users. Guidance is provided to access publicly available PM2.5 datasets, to explain and compare different approaches for dataset generation, and to identify sources of uncertainties associated with various types of datasets. Three main sources used to create PM2.5 exposure data are ground-based measurements (especially regulatory monitoring), satellite retrievals (especially aerosol optical depth, AOD), and atmospheric chemistry models. We find inconsistencies among several publicly available PM2.5 estimates, highlighting uncertainties in the exposure datasets that are often overlooked in health effects analyses. Major differences among PM2.5 estimates emerge from the choice of data (ground-based, satellite, and/or model), the spatiotemporal resolutions, and the algorithms used to fuse data sources.Implications: Fine particulate matter (PM2.5) has large impacts on human morbidity and mortality. Even though the methods for generating the PM2.5 exposure estimates have been significantly improved in recent years, there is a lack of review articles that document PM2.5 exposure datasets that are publicly available and easily accessible by the health and air quality communities. In this article, we discuss the main methods that generate PM2.5 data, compare several publicly available datasets, and show the applications of various data fusion approaches. Guidance to access and critique these datasets are provided for stakeholders in public health sectors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.050 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".