Adjusted Daily Rainfall and Snowfall Data for Canada
Bibliographic record
Abstract
This article documents how Environment and Climate Change Canada’s Adjusted Daily Rainfall and Snowfall (AdjDlyRS) dataset was developed. The adjustments include (i) conversion of ruler measurements of snowfall to its water equivalent using a previously developed snow water equivalent (SWE) ratio map for Canada; (ii) corrections for gauge-related issues including undercatch and evaporation caused by wind effects and gauge-specific wetting loss, as well as for trace precipitation amounts, using previously developed procedures for Canada. Various data flags (e.g., accumulation flags) were also treated. This dataset contains all Canadian stations reporting daily rainfall and snowfall for which we have metadata to implement the adjustments. The length of the data record varies from one station to another, starting as early as 1840. The results show that the original unadjusted total precipitation data in Environment and Climate Change Canada’s digital archive underestimate the total precipitation in northeastern Canada by more than 25% and by about 10–15% in most of southern Canada. Such large underestimates make the original data unsuitable for water availability and/or balance studies or for numerical model validation, among many other applications. The use of the assumed 10:1 SWE ratio for the archived total precipitation data is the primary cause of the underestimate, which is most severe in northeastern Canada. The trace correction adds 5–20% to precipitation values in northern Canada but less than 5% in southern Canada. The gauge-related corrections do not show an organized spatial pattern but add 5–10% to the precipitation at 312 stations. Long runs (≥3 months) of miscoded missing values were also identified and corrected.The latest version of the AdjDlyRS dataset is available from the Canadian Open Data Portal; currently it is version 2016, which contains 3346 stations and covers the period from station inception to February 2016. This dataset is suitable for producing gridded precipitation datasets, as well as other applications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.010 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".