Extending Boundaries of Emission and Dispersion Modelling with Uncertainty Analysis and Data-Driven Models
Bibliographic record
Abstract
Urban transportation systems are undergoing revolutionary changes, including vehicle automation and electrification, and their environmental impacts remain uncertain. Can existing emission and air quality modelling techniques adequately perform in the face of these drastic changes? How can we extend the capabilities of current physical emission models and empirical air quality models? This thesis addresses these questions focusing on two themes: 1) regional traffic emission inventories; 2) data-driven models for local emission and air quality characterization.In theme 1, we first developed a modelling approach to generate a regional emission inventory of passenger transportation (public transit and private household vehicle trips) in the Greater Toronto and Hamilton Area (GTHA). When including only public transit and private household vehicle trips, the latter contributed 96% of regional GHG emissions for passenger transportation. With the introduction of AVs, additional private household vehicle kilometres were expected. Moreover, the GHG emission reduction potential of EVs was highly dependent on the GHG emissions of the local power grid. We identified various sources of uncertainty in current vehicular emission models and quantified these sources of uncertainty with a Monte-Carlo approach. Vehicle operating emissions dominated the total GHG emissions and dominated the uncertainty from other sources. The introduction of EVs also reduced emission uncertainty. Theme 2 focuses on local scale emission and air quality modelling. Noticing the importance of uncertainty in operating emissions, we proposed a novel modal emission modelling approach for vehicle hot stabilized running emissions, based on measured emissions. Our approach can effectively balance computational complexity and accuracy by tuning the resolution of explanatory variables. The last part of this thesis explored the use of a traditional empirical air quality model compared to novel machine learning techniques. Both empirical models were applied to air quality data collected during a mobile sampling campaign in downtown Toronto. We explored their application boundaries by contrasting how explanatory variables were expressed in these data-driven models. We concluded that machine learning could perform better with abundant data, and the prediction power of the traditional empirical model was dependent on the monotonic relationship between explanatory and response variables. 城市交通系统的发展日新月异。首当其冲的是车辆的自动化和电动化。然而相应的环境和能源影响尚未被人们所了解。现行的汽车尾气排放和扩散模型是否能充分捕捉到交通系统变化带来的影响?我们如何能够拓展这些模型的应用范围?本文主要从以下两个主题入手解决这些问题:1. 区域交通排放清单的建立;2. 数据驱动模型在本地尾气排放和扩散模型中的应用。在第一个主题中,我们首先设计了大多伦多和汉密尔顿地区内的客运排放清单的计量方法(包括公共交通和私人交通)。我们发现私人交通产生的温室气体排放占到了客运交通排放的96%。自动驾驶车辆的引入导致了更多的私人交通出行里程。同时,我们观察到电动车带来的温室气体排放减少受到本地电网清洁性的影响。我们进一步确定了数个在现行车辆排放清单模型中的不确定性因素,并将它们用蒙特卡洛方法进行量化。车辆运行中的排放不但占据了区域内燃油生命周期排放的主体,还对区域内车辆排放清单的不确定性产生了主要影响。推行电动车可以减少区域车辆排放清单中的不确定性。 主题二主要探讨了多伦多本地背景下的车辆排放和扩散模型。注意到运行中排放对于区域排放总量和不确定性的影响,我们对此提出了一种新型的排放模型,并用本地采集的排放数据进行开发。我们的方法可以根据具体应用条件来有效平衡计算复杂度和模型准确性。这主要是通过调节解释变量的解析度来实现的。本文的最后一部分探讨了传统的污染物扩散模型(土地利用回归)和机器学习经验模型的比较。我们使用了在多伦多本地采集的空气质量数据来训练这两类模型。通过揭示自变量(空气质量)和因变量(土地利用,交通,天气等)在不同模型中是如何进行表达的,我们探索了不同模型的应用范围。我们发现机器学习模型在有充足数据的情况下可以达到较高的准确度,传统回归模型的鲁棒性需要依靠自变量和因变量之间的单调性关系来保证。
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".