Principal component analysis of summertime ground site measurements in the Athabasca oil sands with a focus on analytically unresolved intermediate-volatility organic compounds
Bibliographic record
Abstract
In this paper, measurements of air pollutants made at a ground site near Fort McKay in the Athabasca oil sands region as part of a multi-platform campaign in the summer of 2013 are presented. The observations included measurements of selected volatile organic compounds (VOCs) by a gas chromatograph–ion trap mass spectrometer (GC-ITMS). This instrument observed a large, analytically unresolved hydrocarbon peak (with a retention index between 1100 and 1700) associated with intermediate-volatility organic compounds (IVOCs). However, the activities or processes that contribute to the release of these IVOCs in the oil sands region remain unclear. Principal component analysis (PCA) with varimax rotation was applied to elucidate major source types impacting the sampling site in the summer of 2013. The analysis included 28 variables, including concentrations of total odd nitrogen (NO y ), carbon dioxide (CO 2 ), methane (CH 4 ), ammonia (NH 3 ), carbon monoxide (CO), sulfur dioxide (SO 2 ), total reduced-sulfur compounds (TRSs), speciated monoterpenes (including α - and β -pinene and limonene), particle volume calculated from measured size distributions of particles less than 10 and 1 µm in diameter (PM 10−1 and PM 1 ), particle-surface-bound polycyclic aromatic hydrocarbons (pPAHs), and aerosol mass spectrometer composition measurements, including refractory black carbon (rBC) and organic aerosol components. The PCA was complemented by bivariate polar plots showing the joint wind speed and direction dependence of air pollutant concentrations to illustrate the spatial distribution of sources in the area. Using the 95 % cumulative percentage of variance criterion, 10 components were identified and categorized by source type. These included emissions by wet tailing ponds, vegetation, open pit mining operations, upgrader facilities, and surface dust. Three components correlated with IVOCs, with the largest associated with surface mining and likely caused by the unearthing and processing of raw bitumen.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".