A benchmark for probabilistic seasonal streamflow forecasting over North America
Bibliographic record
Abstract
Seasonal streamflow forecasts represent critical operational inputs for water sectors and society, for instance for spring flood early warning, water supply, hydropower generation, and irrigation scheduling. Initial hydrological conditions (e.g., snow cover and soil moisture) are an important driver of hydrological predictions on these timescales. In high-latitude and/or high-altitude basins across North America, and the basins downstream of these headwaters, snow is one of the main sources of runoff generation. As a result, data-driven forecasting from snow observations is a well-established approach for operational seasonal streamflow forecasting in the USA (Fleming et al., 2021) and Canada (Zahmatkesh et al., 2019). As part of the Global Water Futures programme (GWF), we are advancing capabilities for probabilistic streamflow forecasting over North America. The first aim of this work is to benchmark probabilistic seasonal streamflow predictability across the continent. To this end, a data-driven probabilistic seasonal streamflow hindcasting system is being developed and implemented for basins with a nival regime across North America. It uses snow water equivalent measurements from the recent update of the Canadian historical Snow Water Equivalent dataset (CanSWE, 1928–2020; Vionnet et al., 2021) and the Natural Resources Conservation Service (NRCS) manual snow surveys and the SNOTEL automatic snow pillow in the USA. These datasets are gap filled using quantile mapping based on neighbouring snow and precipitation stations (SCDNA dataset; Tang et al., 2020), and subsequently transformed into principal components. These principal components are then used as predictors into a regression model, to generate ensemble hindcasts of streamflow volumes for basins across North America. Preliminary results indicate that this approach is skilful (i.e., better than streamflow climatology) for basins across the Canadian Rockies during the snowmelt season. References Fleming, S. W., Garen, D. C., Goodbody, A. G., McCarthy, C. S., and Landers, L. C.: Assessing the new Natural Resources Conservation Service water supply forecast model for the American West: A challenging test of explainable, automated, ensemble artificial intelligence. Journal of Hydrology, 602, https://doi.org/10.1016/j.jhydrol.2021.126782, 2021. Tang, G., Clark, M. P., Newman, A. J., Wood, A. W., Papalexiou, S. M., Vionnet, V., and Whitfield, P. H.: SCDNA: a serially complete precipitation and temperature dataset for North America from 1979 to 2018, Earth Syst. Sci. Data, 12, 2381–2409, https://doi.org/10.5194/essd-12-2381-2020, 2020. Vionnet, V., Mortimer, C., Brady, M., Arnal, L., and Brown, R.: Canadian historical Snow Water Equivalent dataset (CanSWE, 1928–2020), Earth Syst. Sci. Data, 13, 4603–4619, https://doi.org/10.5194/essd-13-4603-2021, 2021. Zahmatkesh, Z., Sanjeev Kumar, J., Coulibaly, P., and Stadnyk, T.: An overview of river flood forecasting procedures in Canadian watersheds, Canadian Water Resources Journal / Revue canadienne des ressources hydriques, 44, 3, https://doi.org/10.1080/07011784.2019.1601598, 2019.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".