A comparison of four receptor models used to quantify the boreal wildfire smoke contribution to surface PM <sub>2.5</sub> in Halifax, Nova Scotia during the BORTAS-B experiment
Bibliographic record
Abstract
Abstract. This paper presents a quantitative comparison of the four most commonly used receptor models, namely absolute principal component scores (APCS), pragmatic mass closure (PMC), chemical mass balance (CMB) and positive matrix factorization (PMF). The models were used to predict the contributions of a wide variety of sources to PM2.5 mass in Halifax, Nova Scotia during the experiment to quantify the impact of BOReal forest fires on Tropospheric oxidants over the Atlantic using Aircraft and Satellites (BORTAS). However, particular emphasis was placed on the capacity of the models to predict the boreal wildfire smoke contributions during the BORTAS experiment. The performance of the four receptor models was assessed on their ability to predict the observed PM2.5 with an R2 close to 1, an intercept close to zero, a low bias and low RSME. Using PMF, a new woodsmoke enrichment factor of 52 was estimated for use in the PMC receptor model. The results indicate that the APCS and PMC receptor models were not able to accurately resolve total PM2.5 mass concentrations below 2 μg m−3. CMB was better able to resolve these low PM2.5 concentrations, but it could not be run on 9 of the 45 days of PM2.5 samples. PMF was found to be the most robust of the four models since it was able to resolve PM2.5 mass below 2 μg m−3, predict PM2.5 mass on all 45 days and utilise an unambiguous woodsmoke chemical tracer. The median woodsmoke relative contributions to PM2.5 estimated using PMC, APCS, CMB and PMF were found to be 0.08, 0.09, 3.59 and 0.14 μg m−3 respectively. The contribution predicted by the CMB model seemed to be clearly too high based on other observations. The use of levoglucosan as a tracer for woodsmoke was found to be vital for identifying this source.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".