A Deep Learning Approach for Meter-Scale Air Quality Estimation in Urban Environments Using Very High-Spatial-Resolution Satellite Imagery
Bibliographic record
Abstract
High-spatial-resolution air quality (AQ) mapping is important for identifying pollution sources to facilitate local action. Some of the most populated cities in the world are not equipped with the infrastructure required to monitor AQ levels on the ground and must rely on other sources, such as satellite derived estimates, to monitor AQ. Current satellite-data-based models provide AQ mapping on a kilometer scale at best. In this study, we focus on producing hundred-meter-scale AQ maps for urban environments in developed cities. We examined the feasibility of an image-based object-detection analysis approach using very high-spatial-resolution (2.5 m) commercial satellite imagery. We fed the satellite imagery to a deep neural network (DNN) to learn the association between visual urban features and air pollutants. The developed model, which solely uses satellite imagery, was tested and evaluated using both ground monitoring observations and land-use regression modeled PM2.5 and NO2 concentrations over London, Vancouver (BC), Los Angeles, and New York City. The results demonstrate a low error with a total RMSE < 2 µg/m3 and highlight the contribution of specific urban features, such as green areas and roads, to continuous hundred-meter-scale AQ estimations. This approach offers promise for scaling to global applications in developed and developing urban environments. Further analysis on domain transferability will enable application of a parsimonious model based merely on satellite images to create hundred-meter-scale AQ maps in developing cities, where current and historical ground data are limited.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".