Predicting outdoor ultrafine particle number concentrations, particle size, and noise using street-level images and audio data
Bibliographic record
Abstract
Outdoor ultrafine particles (UFPs) (<0.1 µm) may have an important impact on public health but exposure assessment remains a challenge in epidemiological studies. We developed a novel method of estimating spatiotemporal variations in outdoor UFP number concentrations and particle diameters using street-level images and audio data in Montreal, Canada. As a secondary aim, we also developed models for noise. Convolutional neural networks were first trained to predict 10-second average UFP/noise parameters using a large database of images and audio spectrogram data paired with measurements collected between April 2019 and February 2020. Final multivariable linear regression and generalized additive models were developed to predict 5-minute average UFP/noise parameters including covariates from deep learning models based on image and audio data along with outdoor temperature and wind speed. The best performing final models had mean cross-validation R2 values of 0.677 and 0.523 for UFP number concentrations and 0.825 and 0.735 for UFP size using two different test sets. Audio predictions from deep learning models were stronger predictors of spatiotemporal variations in UFP parameters than predictions based on street-level images; this was not explained only by noise levels captured in the audio signal. All final noise models had R2 values above 0.90. Collectively, our findings suggest that street-level images and audio data can be used to estimate spatiotemporal variations in outdoor UFPs and noise. This approach may be useful in developing exposure models over broad spatial scales and such models can be regularly updated to expand generalizability as more measurements become available.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".