Chemometric analysis of full scan direct mass spectrometry data for the discrimination and source apportionment of atmospheric volatile organic compounds measured from a moving vehicle.
Bibliographic record
Abstract
Anthropogenic emissions into the troposphere can impact air quality, leading to poorer health outcomes in the affected areas. Volatile organic compounds (VOCs) are a group of chemical compounds, including some which are toxic, that are precursors in the formation of ground-level ozone and secondary organic aerosols. VOCs have a variety of sources, and the distribution of atmospheric VOCs differs significantly over time and space. Historically, the large number of chemical species present at low concentrations (parts-per-trillion to parts-per-billion by volume) have made VOCs difficult to measure in ambient air. However, with improvements in analytical instrumentation, these measurements are becoming more common place. Direct mass spectrometry (MS), such as membrane introduction mass spectrometry (MIMS) and proton-transfer reaction time-of-flight mass spectrometry (PTR-ToF-MS) facilitate real-time, continuous measurements of VOCs in air, with full scan mass spectral data capturing changes in chemical composition with high temporal resolution. Operated on-road, mobilized direct MS has been used for quantitative mapping of VOCs at the neighborhood scale, but identifying VOC sources based on the observed mixture of molecules in the full scan MS dataset has yet to be explored. This dissertation describes the use of chemometric techniques to interrogate full scan MS data, and the progression from discriminating VOC samples of known chemical composition based on full scan MIMS data through to the apportionment of VOC sources measured continuously with a PTR-ToF-MS system operating in a moving vehicle. Lab‐constructed VOC samples of known chemical composition and concentration demonstrated the use of principal component analysis (PCA) to discriminate, and k-nearest neighbours to classify, samples based on normalized full scan MIMS data. Furthermore, multivariate curve resolution-alternating least squares (MCR-ALS) was used to resolve mixtures into molecular component contributions. PCA was also used to discriminate ‘real-world’ VOC mixtures (e.g., woodsmoke VOCs, headspace above aqueous hydrocarbon samples) of unknown chemical composition measured by MIMS. Using vehicle mounted MIMS and PTR-ToF-MS systems, full scan MS data of ambient atmospheric VOCs were collected and PCA was applied to the normalized full scan MS data. A supervised analysis performed PCA on samples collected near known VOC sources, while an unsupervised analysis using PCA followed by cluster analysis was used to identify groups in a continuous, time series PTR-ToF-MS dataset measured between Nanaimo and Crofton, British Columbia (BC). In both the supervised and unsupervised analysis, samples impacted by emissions from different sources (e.g., internal combustion engines, sawmills, composting facilities, pulp mills) were discriminated. With PCA, samples were discriminated based on differences in the observed full scan MS data, however real-world samples are often impacted by multiple VOC sources. MCR-weighted ALS (MCR-WALS) was applied to the continuous, time series PTR-ToF-MS data from three field campaigns on Vancouver Island, BC for source apportionment. Variable selection based on signal-to-noise ratios was used to reduce the mass list while retaining the observed m/z that capture changes in the mixture of VOCs measured, improving model results, and reducing computation time. Both point (e.g., anthropogenic hydrocarbon emissions, pulp mill emissions) and diffuse (e.g., VOCs from forest fire smoke) VOC sources were identified in the data, and were apportioned to determine their contributions to the measured samples. The data analyzed captured fine scale changes in the ambient VOCs present in the air, and geospatial maps of each individual source, and of the source apportionment were used to visualize the distribution of VOC sources across the sampling area. This work represents the first use of MCR-WALS to identify and apportion ambient VOC sources based on continuous PTR-ToF-MS data measured from a moving vehicle. The methods described can be applied to larger scale field campaigns for the source apportionment of VOCs across multiple days to capture diurnal and seasonal variations. Identifying spatial and temporal trends in the sources of VOCs at the regional scale can help to identify pollution ‘hot spots’ and inform evidence-based public policy for improving air quality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".