MétaCan
Menu
← Back to cohort
Record W4221006948 · doi:10.5194/egusphere-egu22-6280

A benchmark for probabilistic seasonal streamflow forecasting over North America

2022· preprint· en· W4221006948 on OpenAlexaffabout
Louise Arnal, Martyn Clark, Vincent Fortin, Alain Pietroniro, Vincent Vionnet, Paul H. Whitfield, Andrew W. Wood

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldEarth and Planetary Sciences
TopicCryospheric studies and observations
Canadian institutionsEnvironment and Climate Change CanadaUniversity of CalgaryUniversity of Saskatchewan
Fundersnot available
KeywordsStreamflowEnvironmental scienceFlood forecastingSnowClimatologySnowmeltSurface runoffPrecipitationHydropowerHydrology (agriculture)MeteorologyDrainage basinGeographyGeology

Abstract

fetched live from OpenAlex

Seasonal streamflow forecasts represent critical operational inputs for water sectors and society, for instance for spring flood early warning, water supply, hydropower generation, and irrigation scheduling. Initial hydrological conditions (e.g., snow cover and soil moisture) are an important driver of hydrological predictions on these timescales. In high-latitude and/or high-altitude basins across North America, and the basins downstream of these headwaters, snow is one of the main sources of runoff generation. As a result, data-driven forecasting from snow observations is a well-established approach for operational seasonal streamflow forecasting in the USA (Fleming et al., 2021) and Canada (Zahmatkesh et al., 2019). As part of the Global Water Futures programme (GWF), we are advancing capabilities for probabilistic streamflow forecasting over North America. The first aim of this work is to benchmark probabilistic seasonal streamflow predictability across the continent. To this end, a data-driven probabilistic seasonal streamflow hindcasting system is being developed and implemented for basins with a nival regime across North America. It uses snow water equivalent measurements from the recent update of the Canadian historical Snow Water Equivalent dataset (CanSWE, 1928–2020; Vionnet et al., 2021) and the Natural Resources Conservation Service (NRCS) manual snow surveys and the SNOTEL automatic snow pillow in the USA. These datasets are gap filled using quantile mapping based on neighbouring snow and precipitation stations (SCDNA dataset; Tang et al., 2020), and subsequently transformed into principal components. These principal components are then used as predictors into a regression model, to generate ensemble hindcasts of streamflow volumes for basins across North America. Preliminary results indicate that this approach is skilful (i.e., better than streamflow climatology) for basins across the Canadian Rockies during the snowmelt season. References Fleming, S. W., Garen, D. C., Goodbody, A. G., McCarthy, C. S., and Landers, L. C.: Assessing the new Natural Resources Conservation Service water supply forecast model for the American West: A challenging test of explainable, automated, ensemble artificial intelligence. Journal of Hydrology, 602, https://doi.org/10.1016/j.jhydrol.2021.126782, 2021. Tang, G., Clark, M. P., Newman, A. J., Wood, A. W., Papalexiou, S. M., Vionnet, V., and Whitfield, P. H.: SCDNA: a serially complete precipitation and temperature dataset for North America from 1979 to 2018, Earth Syst. Sci. Data, 12, 2381–2409, https://doi.org/10.5194/essd-12-2381-2020, 2020. Vionnet, V., Mortimer, C., Brady, M., Arnal, L., and Brown, R.: Canadian historical Snow Water Equivalent dataset (CanSWE, 1928–2020), Earth Syst. Sci. Data, 13, 4603–4619, https://doi.org/10.5194/essd-13-4603-2021, 2021. Zahmatkesh, Z., Sanjeev Kumar, J., Coulibaly, P., and Stadnyk, T.: An overview of river flood forecasting procedures in Canadian watersheds, Canadian Water Resources Journal / Revue canadienne des ressources hydriques, 44, 3, https://doi.org/10.1080/07011784.2019.1601598, 2019.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.006
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.863
Threshold uncertainty score0.272

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.006
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.002
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.044
GPT teacher head0.238
Teacher spread0.194 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes2
Has abstractyes

Explore more

Same topicCryospheric studies and observations→French-language works237,207→