Traffic NO<sub>x</sub> Pollution Prediction and Health Cost Estimation Using Machine Learning: A Case Study of Toronto, Canada
Bibliographic record
Abstract
Road traffic is a significant source of air pollution that has a harmful impact on human health. To reduce the health and environmental impacts of fossil fuel consumption in the transportation sector, many countries have implemented policies to promote the deployment of electric vehicles (EVs). A vital factor to consider when designing policies to support EV use is the monetized health impacts of fossil fuel consumption. This research aims to investigate the health benefit of replacing internal combustion engine vehicles (ICEVs) with zero-emission vehicles in the city of Toronto, Canada. A long short-term memory (LSTM) model is developed in this work to predict future NOx concentrations considering the effect of the traffic volume, weather, time of day, and historical NOx concentrations. The developed model is then used to predict long-term NOx concentrations and annual average NOx reduction from zero-emission vehicle deployment in four different scenarios in Toronto. Additionally, interpolation methods are used to predict the pollution reduction in all Dissemination Areas (DA) of Toronto, and a health cost assessment is conducted to estimate the health benefit in all the scenarios. The results of the modeling in this work show that the western areas of Toronto experience higher NOx concentration reduction in all scenarios. These reductions are the result of the higher correlation between traffic volume and pollution in those areas. The results also show that with a 10% reduction in ICEV traffic volume, 70 premature deaths can be prevented annually, equivalent to 560 million CAD in health benefits per year.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".