Machine Learning Classification of Air Quality Monitoring Stations to Achieve Ambient NO2 Objectives Using Emission Scenarios and Chemical Transport Model
Bibliographic record
Abstract
Exceeding the latest Canadian Ambient Air Quality Standards (CAAQS) for NO 2 concentration in Canadian cities motivates discussions on new policies and directives. Advanced air quality modeling and machine learning clustering algorithms integrating the Weather Research Forecast (WRF) and Community Multiscale Air Quality (CMAQ) models were used to classify air quality monitoring (AQM) stations based on pollution concentration modeling results. The sensitivity of ambient NO 2 to the primary anthropogenic emission sources was investigated in Alberta, Canada. Emissions from two main sources, upstream oil and gas (UOG) and transportation sources, have been identified as the reason Alberta fails to meet the newly adopted NO 2 CAAQS objectives. The air quality model was validated with ground-level observation data and the atmospheric model accurately replicates spatiotemporal NO 2 variations. Despite contributing 62% of Alberta’s total NO x emissions, UOG influences ambient NO 2 concentrations modestly in urban areas ( < 10%) but significantly affects rural regions. In contrast, transportation emission sources, responsible for 23% of NO x emissions, dominate ambient NO 2 levels (up to 63%) in large cities. The discrepancy of emission contribution and ground-level concentrations, obtained from the chemical transport model, was resolved using a k-prototypes clustering algorithm to propose a new approach for categorizing AQM stations which led to improving the conventional classifications. The new approach considered the sensitivity of NO 2 to emission reduction scenarios and provided an improved classification to be used for emission reduction interventions. Based on the updated classification, one set of AQM stations clearly showed sensitivity to NO 2 emission reduction in the transportation sector despite their lower contributions to overall emissions. These stations were categorized as population exposure stations in large cities. • Innovative approach combining Chemical Transport Model and Machine Learning to classify air quality monitoring stations. • Combined consideration of primary source and land use for emission reduction policies. • Brute Force Sensitivity Analysis of emission reduction scenarios on ambient NO 2 concentration. • Due to proximity to population, transportation emissions dominate NO 2 concentration even in the fourth-largest global oil reservoir.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".