Defining a reference climate for canada through the \ncombination of observational and reanalysis datasets.
Bibliographic record
Abstract
Le présent travail propose une combinaison d’ensemble de différents jeux de données à l’aide de \ndeux méthodes d’assimilation pour améliorer la qualité des champs de précipitation climatiques au \nCanada pour une période de 30 ans (1980-2009). Quatre ensembles de données de réanalyse ont \nété utilisés : CFSR, ERAI, JRA55 et MERRA. Les données d’observation proviennent de 2160 \nstations météorologiques à travers le pays avec un minimum de 10 années valides de précipitations \nquotidiennes. De plus, les données d’observation sur grille de Ressources Naturelles Canada \n(NRCan pour Natural Resources Canada) ont été utilisées à fin de comparaison. Les Indices de \nPrécipitation Climatiques (IPC) ont été calculés pour chaque ensemble de données annuellement \net ont été combinés (la Réanalyse + les Observations) par les méthodes d’Interpolation Optimale \n(OI pour Optimal Interpolation) et d’Interpolation Optimale d’Ensemble (EnOI pour Ensemble \nOptimal Interpolation) respectivement. Une vérification a été effectuée sur les stations qui ont \nété utilisées pour la combinaison, de même qu’une stratégie de validation croisée. Les résultats \nont montré des améliorations au niveau des nouveaux ensembles de données de référence créés par \nles deux méthodes, en particulier pour l’approche d’ensemble. L’ensemble de données sur grille \ndéveloppé (par EnOI) a surpassé NRCan, lequel a été généré seulement à partir d’observations, \nce qui montre que ce dernier devrait être utilisé avec précaution pour l’analyse d’événements de \nprécipitation extrême. L’ensemble de données produit pourrait être utilisé pour les études d’impact \nde changement climatique et de stratégies d’adaptation de même que pour fournir une meilleure \ncompréhension du climat actuel au Canada. The present work shows a combination of different datasets through two data assimilation methods \nin order to improve the quality of the climate precipitation fields in Canada for a period of 30 \nyears (1980-2009). Four reanalysis datasets were used; CFSR, ERAI, JRA55 and MERRA. The \nobservational dataset consists in 2160 meteorological stations across the country with a minimum \nof 10 valid years of precipitation records available. Besides, the Natural Resources Canada gridded \ndataset (NRCan) was also used. The Climate Precipitation Indices were calculated for each dataset \nannually and were combined (Reanalysis + Observations) by the Optimal Interpolation (OI) and \nEnsemble Optimal Interpolation (EnOI) methods respectively. The verification was carried out over \nthe stations that were used in the combination, but also through a cross-validation strategy. Results \nshow the improvements in the new reference datasets created by both methods highlighting the \nEnsemble approach as the best. The reference gridded dataset developed (by EnOI) outperformed \nNRCan which was created only by observations and showing that the latter should be used \nwith caution for extreme precipitation events analysis. The dataset produced contributes to the \ndevelopment of climate change impact studies and adaptation strategies, as well as providing a \nbetter understanding of the current climate in Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.006 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".