Machine Learning Classification of Air Quality Monitoring Stations to Achieve Ambient NO2 Objectives Using Emission Scenarios and Chemical Transport Model
Bibliographic record
Abstract
Exceeding the latest Canadian Ambient Air Quality Standards (CAAQS) for NO2 concentration in Canadian cities motivates discussions on new policies and directives. Advanced air quality modeling and machine learning clustering algorithms integrating the Weather Research Forecast (WRF) and Community Multiscale Air Quality (CMAQ) models were used to classify air quality monitoring (AQM) stations based on pollution concentration modeling results. The sensitivity of ambient NO2 to the primary anthropogenic emission sources was investigated in Alberta, Canada. Emissions from two main sources, upstream oil and gas (UOG) and transportation sources, have been identified as the reason Alberta fails to meet the newly adopted NO2 CAAQS objectives. The air quality model was validated with ground-level observation data and the atmospheric model accurately replicates spatiotemporal NO2 variations. Despite contributing 62% of Alberta’s total NOx emissions, UOG influences ambient NO2 concentrations modestly in urban areas (<10%) but significantly affects rural regions. In contrast, transportation emission sources, responsible for 23% of NOx emissions, dominate ambient NO2 levels (up to 63%) in large cities. The discrepancy of emission contribution and ground-level concentrations, obtained from the chemical transport model, was resolved using a k-prototypes clustering algorithm to propose a new approach for categorizing AQM stations which led to improving the conventional classifications. The new approach considered the sensitivity of NO2 to emission reduction scenarios and provided an improved classification to be used for emission reduction interventions. Based on the updated classification, one set of AQM stations clearly showed sensitivity to NO2 emission reduction in the transportation sector despite their lower contributions to overall emissions. These stations were categorized as population exposure stations in large cities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".