MétaCan
Menu
← Retour à la cohorte
Enregistrement W7116840947 · doi:10.5194/egusphere-2025-4324

Operational chemical weather forecasting with the ECCC online Regional Air Quality Deterministic Prediction System version 023 (RAQDPS023) – Part 2: Multi-year prospective and retrospective performance evaluation

2025· article· W7116840947 sur OpenAlexafffundabout
Michael D. Moran, Alexandru Lupu, Verica Savic-Jovcic, Junhua Zhang, Qiong Zheng, Elisa I. Boutzis, Rabab Mashayekhi, Craig Stroud, Sylvain Ménard, J. Chen, Konstantinos Menelaou, Rodrigo Munoz-Alpizar, Dragana Kornic, Patrick M. Manseau

Notice bibliographique

Revuenon disponible
Typearticle
Langue
DomaineEarth and Planetary Sciences
ThématiqueAtmospheric chemistry and aerosols
Établissements canadiensEnvironment and Climate Change Canada
Organismes subventionnairesEnvironment and Climate Change CanadaU.S. Environmental Protection Agency
Mots-clésAir quality indexData setQuality (philosophy)Period (music)Air temperatureWeather forecasting

Résumé

récupéré en direct d'OpenAlex

Abstract. The online version of the Regional Air Quality Deterministic Prediction System (RAQDPS) is a chemical weather forecast system that has been employed operationally by Environment and Climate Change Canada (ECCC) since 2009. It is run twice daily to produce 72 hour forecasts of hourly 10 km abundance fields of three key predictands, NO2, O3, and PM2.5 total mass, as well as other gas-phase chemical species, PM2.5 chemical components, and dry and wet deposition for Canada, the contiguous U.S., and northern Mexico. Version 023 of the RAQDPS (RAQDPS023) went into service at ECCC in December 2021 and was replaced by the RAQDPS025 in June 2024. A companion paper by Moran et al. (2025) describes the RAQDPS023 in detail. In this paper we present the results of a five-year performance evaluation of prospective and retrospective annual air quality (AQ) simulations made with the RAQDPS023. The annual simulations considered were the first year of RAQDPS023 forecasts in 2021/22 and four years of retrospective annual simulations for the 2013‒2016 period that used historical, year-specific emissions. Forecasts made by the RAQDPS-FW023, a duplicate operational system to the RAQDPS023 except for the addition of near-real-time (NRT) biomass burning (BB) emissions, were also evaluated for the 2021/22 period. A NRT measurement data set consisting of hourly NO2, O3, and PM2.5 surface measurements for Canada and the U.S. was used for the 2021/22 evaluation whereas a much more extensive set of air-chemistry and precipitation-chemistry measurements was used for the 2013‒2016 evaluations. Some evaluation results were also compared with results for the 2010‒2019 period for forecasts made by earlier operational versions of the RAQDPS and with evaluation results for several peer AQ forecast models. In addition to looking at a number of highly aggregated “headline” scores, many stratified analyses were also performed, including evaluations by network, season, month, hour of day, region, and land-use type. Consideration of simulations for multiple years with the same model and year-specific input emissions helped to identify systematic model errors by reducing the influence of year-to-year variations in meteorology, and a comprehensive evaluation for many species for 2013‒2016 supported by stratified analyses provided diagnostic insights that allowed the scientific basis for the RAQDPS023 forecasts to be assessed (i.e., “right answers for the right reasons?”). Although one confounding factor for this study was the sizable reduction in the emissions of some pollutants in North America that occurred from 2013 to 2021, it was found that the trends in AQ observations over this period agreed with the year-specific description of emissions used for the five annual simulations from a rank-ordered perspective. While RAQDPS023 evaluation scores for hourly NO2 and O3 volume mixing ratio forecasts were found to be competitive with peer models and often met suggested performance benchmarks for the five simulation years, another key finding was that the RAQDPS023 forecasts consistently underpredicted hourly PM2.5 total mass concentrations for all months in 2021/22 and for the majority of months in 2013‒2016. The largest underpredictions occurred in summer and at rural stations whereas overpredictions often occurred in the cold season at urban stations. The model also missed the observed bimodality in monthly PM2.5 concentrations and exaggerated the observed diurnal variations in hourly PM2.5 concentrations. Additional evaluations with daily PM2.5 chemical composition measurements and daily gravimetric PM2.5 total mass measurements were also examined to better understand the hourly PM2.5 underpredictions. Consistent overpredictions of elemental carbon and sea salt concentrations and underpredictions of sulfate concentration were identified, but scores for predictions of daily gravimetric PM2.5 total mass were better than those for hourly PM2.5 total mass, directing attention to differences in measurement methods. SO2 and HNO3 levels were also found to be overpredicted in general while NH3 levels were underpredicted: these three gas-phase species are all PM2.5 precursors, which raises concerns about some process representations such as those for sulfur oxidation and gas-phase dry deposition. As well, springtime O3 levels were underpredicted while isoprene levels were consistently overpredicted. The impact of BB emissions on predictions of NO2, O3, and PM2.5 was also characterized in detail by comparing evaluation results for the 2021/22 RAQDPS023 and RAQDPS-FW023 forecasts. Negligible impact was found for monthly NO2 forecasts when BB emissions were included, but monthly O3 forecast scores were modestly improved and monthly PM2.5 forecast scores were markedly improved from July to September 2021, as well as summer and annual scores. Taken together, the results of this comprehensive multi-year evaluation point to a number of RAQDPS023 system components where improvements are desirable. These results also provide a strong benchmark against which to compare the performance of future versions of the RAQDPS.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,004
score de la tête « metaresearch » (Gemma)0,006
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,051
Score d'incertitude au seuil0,101

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0040,006
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0000,001
Études des sciences et des technologies0,0000,000
Communication savante0,0010,001
Science ouverte0,0010,000
Intégrité de la recherche0,0010,001
Charge utile insuffisante (le modèle a refusé de juger)0,0010,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,033
Tête enseignante GPT0,255
Écart entre enseignants0,222 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2025
Routes d'admission3
Résumé présentoui

Explorer davantage

Même sujetAtmospheric chemistry and aerosols→Travaux en français237 207→