Reliable probabilistic forecasts from an ensemble reservoir inflow forecasting system
Bibliographic record
Abstract
Abstract This paper describes a probabilistic reservoir inflow forecasting system that explicitly attempts to sample from major sources of uncertainty in the modeling chain. Uncertainty in hydrologic forecasts arises due to errors in the hydrologic models themselves, their parameterizations, and in the initial and boundary conditions (e.g., meteorological observations or forecasts) used to drive the forecasts. The Member‐to‐Member (M2M) ensemble presented herein uses individual members of a numerical weather model ensemble to drive two different distributed hydrologic models, each of which is calibrated using three different objective functions. An ensemble of deterministic hydrologic states is generated by spinning up the daily simulated state using each model and parameterization. To produce probabilistic forecasts, uncertainty models are used to fit probability distribution functions (PDF) to the bias‐corrected ensemble. The parameters of the distribution are estimated based on statistical properties of the ensemble and past verifying observations. The uncertainty model is able to produce reliable probability forecasts by matching the shape of the PDF to the shape of the empirical distribution of forecast errors. This shape is found to vary seasonally in the case‐study watershed. We present an “intelligent” adaptation to a Probability Integral Transform (PIT)‐based probability calibration scheme that relabels raw cumulative probabilities into calibrated cumulative probabilities based on recent past forecast performance. As expected, the intelligent scheme, which applies calibration corrections only when probability forecasts are deemed sufficiently unreliable, improves reliability without the inflation of ignorance exhibited in certain cases by the original PIT‐based scheme.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".