Evaluation of several PM<sub>2.5</sub> forecast models using data collected during the ICARTT/NEAQS 2004 field study
Bibliographic record
Abstract
Real‐time forecasts of PM2.5 aerosol mass from seven air quality forecast models (AQFMs) are statistically evaluated against observations collected in the northeastern United States and southeastern Canada from two surface networks and aircraft data during the summer of 2004 International Consortium for Atmospheric Research on Transport and Transformation (ICARTT)/New England Air Quality Study (NEAQS) field campaign. The AIRNOW surface network is used to evaluate PM2.5 aerosol mass, the U.S. EPA STN network is used for PM2.5 aerosol composition comparisons, and aerosol size distribution and composition measured from the NOAA P‐3 aircraft are also compared. Statistics based on midday 8‐hour averages, as well as 24‐hour averages are evaluated against the AIRNOW surface network. When the 8‐hour average PM2.5 statistics are compared against equivalent ozone statistics for each model, the analysis shows that PM2.5 forecasts possess nearly equivalent correlation, less bias, and better skill relative to the corresponding ozone forecasts. An analysis of the diurnal variability shows that most models do not reproduce the observed diurnal cycle at urban and suburban monitor locations, particularly during the nighttime to early morning transition. While observations show median rural PM2.5 levels similar to urban and suburban values, the models display noticeably smaller rural/urban PM2.5 ratios. The ensemble PM2.5 forecast, created by combining six separate forecasts with equal weighting, is also evaluated and shown to yield the best possible forecast in terms of the statistical measures considered. The comparisons of PM2.5 composition with NOAA P‐3 aircraft data reveals two important features: (1) The organic component of PM2.5 is significantly underpredicted by all the AQFMs and (2) those models that include aqueous phase oxidation of SO2 to sulfate in clouds overpredict sulfate levels while those AQFMs that do not include this transformation mechanism underpredict sulfate. Errors in PM2.5 ammonium levels tend to correlate directly with errors in sulfate. Comparisons of PM2.5 composition with the U.S. EPA STN network for three of the AQFMs show that sulfate biases are consistently lower at the surface than aloft. Recommendations for further research and analysis to help improve PM2.5 forecasts are also provided.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".