Comparing different versions of the continuous ranked probability score to account for forecast or observation uncertainty
Bibliographic record
Abstract
Recent studies have shown that probabilistic forecasts are superior to deterministic forecasts in terms of quality, reliability, and representing the uncertainty of future states. One of the most well-known and widely used tools for assessing the performance of (probabilistic) forecast systems is the continuous ranked probability score (CRPS). This metric is employed to evaluate the forecasting system when only forecast uncertainty is considered. In addition to multiple sources of uncertainty in a forecasting system (such as initial conditions, model structure and parameters, and boundary conditions), the uncertainty can also originate from observations (e.g., streamflow). However, this uncertainty, which has rarely been explored in previous research, should also be regarded in evaluating the forecasting system. A version of the CPRS is redefined and analyzed to overcome this important flaw, considering the observation's uncertainty. To estimate the uncertainty associated with streamflow observations, the Bayesian Rating curve method (BaRatin) is utilized. This study focuses on comparing the different versions of the CRPS in considering the uncertainties of forecasts and observations. Three types of streamflow forecasting systems are used in this study: deterministic forecasts, raw ensemble forecasts (applying meteorological ensemble forecasts as inputs to the hydrological model), and post-processed ensemble forecasts (postprocessing of hydrological model outputs using weighted ensemble dressing method). The assessment is performed for short-term forecasts (lead times of 1 to 5 days) for the Au Saumon watershed in southern central Quebec, Canada. It is found that considering observation uncertainty has a significant effect on the values of CRPS compared to when only forecast uncertainty is considered. In addition, CRPS changes in probabilistic forecasts are more than deterministic ones. Our results also point out that using the modified version of the CRPS can help end-users better understand and evaluate their forecasting system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.064 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".