Identifying surface sulphur dioxide (SO2) monitoring gaps in Saint John, Canada with land use regression and hot spot mapping
Bibliographic record
Abstract
Saint John experiences ambient sulphur dioxide (SO 2 ) pollution due to a high density of industrial activities. Despite recent reduction in SO 2 emissions, over 90 % of the provincial exceedances of air pollutants were related to SO 2 or total reduced sulphur (TRS), and over 70 % among which occurred in Saint John. Pinpointing intra-urban SO 2 hot spots is important for revealing the neighborhoods exposed to high health risk. However, this is challenging due to limited spatial coverage of monitoring. To fill the monitoring gap, we developed two-stage gradient boosting models combining a classifier that discerned between SO 2 -free and SO 2 -polluted days and a regressor that estimated daily SO 2 levels based on remote sensing data. With a 10-fold cross-validation, the classifier achieved 83 % accuracy and the regressors attained R 2 of 0.46 and 0.44 for daily mean and maximum SO 2 respectively. Based on model outputs, we conducted spatial hot spot analysis and found high SO 2 levels spread to northeast, north, and southeast Saint John, where SO 2 monitoring was absent. Several existing monitoring sites in west Saint John do not have SO 2 regularly measured. Besides the spatiotemporal lags of nearby monitored SO 2 , wind-related variables such as wind speed and direction had high importance in predicting surface SO 2 , which might suggest potential impacts to remote unmonitored communities from the transport of SO 2 . In summary, our findings suggest that certain unmonitored areas in Saint John may experience high SO 2 levels. Expansion of monitoring efforts would help inform where and when mitigation should be taken to minimize SO 2 -related health impacts. • We employed a two-stage modelling approach for estimating surface SO 2 in Saint John. • Model-derived SO 2 hot spots helped identify areas for additional air monitoring. • Increased SO 2 monitoring is needed in the northern and eastern parts of Saint John. • Wind-related variables and temporal lags of surface SO 2 greatly impacted the result.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".