Gas and aerosol carbon in California: comparison of measurements and model predictions in Pasadena and Bakersfield
Bibliographic record
Abstract
Abstract. Co-located measurements of fine particulate matter (PM2.5) organic carbon (OC), elemental carbon, radiocarbon (14C), speciated volatile organic compounds (VOCs), and OH radicals during the CalNex field campaign provide a unique opportunity to evaluate the Community Multiscale Air Quality (CMAQ) model's representation of organic species from VOCs to particles. Episode average daily 23 h average 14C analysis indicates PM2.5 carbon at Pasadena and Bakersfield during the CalNex field campaign was evenly split between contemporary and fossil origins. CMAQ predicts a higher contemporary carbon fraction than indicated by the 14C analysis at both locations. The model underestimates measured PM2.5 organic carbon at both sites with very little (7% in Pasadena) of the modeled mass represented by secondary production, which contrasts with the ambient-based SOC / OC fraction of 63% at Pasadena. Measurements and predictions of gas-phase anthropogenic species, such as toluene and xylenes, are generally within a factor of 2, but the corresponding SOC tracer (2,3-dihydroxy-4-oxo-pentanoic acid) is systematically underpredicted by more than a factor of 2. Monoterpene VOCs and SOCs are underestimated at both sites. Isoprene is underestimated at Pasadena and overpredicted at Bakersfield and isoprene SOC mass is underestimated at both sites. Systematic model underestimates in SOC mass coupled with reasonable skill (typically within a factor of 2) in predicting hydroxyl radical and VOC gas-phase precursors suggest error(s) in the parameterization of semivolatile gases to form SOC. Yield values (α) applied to semivolatile partitioning species were increased by a factor of 4 in CMAQ for a sensitivity simulation, taking into account recent findings of underestimated yields in chamber experiments due to gas wall losses. This sensitivity resulted in improved model performance for PM2.5 organic carbon at both field study locations and at routine monitor network sites in California. Modeled percent secondary contribution (22% at Pasadena) becomes closer to ambient-based estimates but still contains a higher primary fraction than observed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".