A comparative study of two factor analytic models applied to PAH data from inhalable air particulate collected in an urban-industrial environment
Bibliographic record
Abstract
Two factor analysis (FA)-based receptor modeling methods were applied to a polycyclic aromatic hydrocarbon (PAH) dataset from extracts of 75 PM(10) air particulate samples collected concurrently at 4 sampling sites proximate to the urban-industrial area in Hamilton, Ontario, Canada. The total PAH concentrations of 48 target compounds ranged from 0.23 to 172 ng m(-3). Principal component analysis (PCA) and positive matrix factorization (PMF) analysis were followed by multilinear regression analyses to identify and quantify PAH source contributions, together with spatial and temporal trends. The correlations between predicted and observed total PAH levels were excellent in both models (R(2) > 0.98). The PCA afforded large negative contributions in a number of samples, so further analysis was abandoned. The PMF analysis showed 3 factors which were identified as gasoline emissions, diesel emissions and coke oven emissions. Contributions of gasoline emissions and diesel emissions factors were surprisingly similar at all 4 sites indicative of a background of vehicle emissions across the city. The PMF coke oven emission factor showed the greatest variability in total loadings, consistent with the large PAH emissions from the steel industries and the large influence of wind direction on PAH concentrations. The highest coke oven contributions were observed at sites closest to the industrial area on days when these sites were downwind of the industries. The PMF coke oven impact factor showed good correlations with two commonly used PAH diagnostic ratios when the ratios were combined into a single ratio. This integrated approach allowed us to categorize >90% of the samples based on the wind direction of the impacting source.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.043 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".