Exploring the relationship between the skill of hydrological ensemble predictions and catchment descriptors
Bibliographic record
Abstract
Skillful hydrological forecasts are essential for decision-making in many areas such as preparedness against natural disasters, water resources management, and hydropower operations. Despite the great technological advances, obtaining skillful predictions from a forecasting system, under a range of conditions and geographic locations, remain a difficult task. It is still unclear why some systems perform better than others at different temporal and spatial scales. Much work has been devoted to investigate the quality of forecasts and the relative contributions of meteorological forcing, catchment’s initial conditions, and hydrological model structure in a streamflow forecasting system. These sources of uncertainty are rarely considered fully and simultaneously in operational systems, and there are still gaps in understanding their relationship with the dominant processes and mechanisms that operate in a given river basin. In this study, we use a multi-model hydrological ensemble prediction system (H-EPS) named HOOPLA (HydrOlOgical Prediction Laboratory), which allows to account separately for these three main sources of uncertainty in hydrological ensemble forecasting. Through the use of EnKF data assimilation, of 20 lumped hydrological models, and of the 50-member ECMWF medium-range weather forecasts, we explore the relationship between the skill of ensemble predictions and the many descriptors (e.g. catchment surface, climatology, morphology, flow threshold and hydrological regime) that influence hydrological predictability. We analyze streamflow forecasts at 50 stations spread across Quebec, France and Colombia, over the period from 2011 to 2015 and for lead times up to 9 days. The forecast performance is assessed using common metrics for forecast quality verification, such as CRPS, Brier skill score, and reliability diagrams. Skill scores are computed using a probabilistic climatology benchmark, which was generated with the hydrological models forced by resampled historical meteorological data. Our results contribute to relevant literature on the topic and bring additional insight into the role of each descriptor in the skill of a hydrometeorological ensemble forecasting chain, serving as a possible guide for potential users to identify the circumstances or conditions in which it is more efficient to implement a given system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.015 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".