Comparative Performance of Different Statistical Models for Predicting Ground-Level Ozone (O3) and Fine Particulate Matter (PM2.5) Concentrations in Montréal, Canada
Bibliographic record
Abstract
Ground-level ozone (O3) and fine particulate matter (PM2.5) are two air pollutants known to reduce visibility, to have damaging effects on building materials and adverse impacts on human health. O3 is the result of a series of complex chemical reactions between nitrogen oxides (NOx) and volatile organic compounds (VOCs) in the presence of solar radiation. PM is a class of airborne contaminants composed of sulphate, nitrate, ammonium, crustal components and trace amounts of microorganisms. PM2.5 is the respirable subgroup of PM having an aerodynamic diameter of less than 2.5 μm. Development of effective forecasting models for ground-level O3 and PM2.5 is important to warn the public about potentially harmful or unhealthy concentration levels. \n \nThe objectives of this study is to investigate the applicability of Multiple Linear Regression (MLR), Principle Component Regression (PCR), Multivariate Adaptive Regression Splines (MARS), feed-forward Artificial Neural Networks (ANN) and hybrid Principal Component – Artificial Neural Networks (PC-ANN) models to predict concentrations of O3 and PM2.5 in Montréal (Canada). Air quality and meteorological data is obtained from the Réseau de surveillance de la qualité de l’air (RSQA) for the Airport Station (45°28′N, 73°44′W) and the Maisonneuve Station (45°30′N, 73°34′W) for the period January 2004 to December 2007. Air pollution data include concentration values for nitrogen monoxide (NO), nitrogen dioxide (NO2), carbon monoxide (CO) and 142 different volatile organic compounds. Meteorological data include solar irradiation (SR), temperature (Temp), pressure (Press), dew point (DP), precipitation (Precip), wind speed (WS) and wind direction (WD). \n \nAnalysis of the available volatile organic compound data expressed on a propylene-equivalent concentration indicated that m/p-xylene, toluene, propylene and (1,2,4)-trimethylbenzene were species with the most significant ozone forming potential in the study area. \n \nDifferent models and architectures have been investigated through five case studies. Predictive performances of each model have been measured by means of performance metrics and forecast success rates. Overall, MARS models allowing second order interaction of independent basis functions yielded lower error, higher correlation and higher forecast success rates. This study indicates that models based on statistical methods can be cost-effective tools to forecast ground-level O3 and PM2.5 in Montréal and to provide support for decision makers in protecting human health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".