On the Choice of Metric to Calibrate Time-Invariant Ensemble Kalman Filter Hyper-Parameters for Discharge Data Assimilation and Its Impact on Discharge Forecast Modelling
Bibliographic record
Abstract
An important step when using some data assimilation methods, such as the ensemble Kalman filter and its variants, is to calibrate its parameters. Also called hyper-parameters, these include the model and observation errors, which have previously been shown to have a strong impact on the performance of the data assimilation method. Many metrics can be used to calibrate these hyper-parameters but may not all yield the same optimal set of values. The current study investigated the importance of the choice of metric used during the hyper-parameter calibration phase and its impact on discharge forecasts. The types of metrics used each focused on discharge accuracy, ensemble spread or observation-minus-background statistics. The calibration was performed for the ensemble square root Kalman filter over two catchments in Canada using two different hydrologic models per catchment. Results show that the optimal set of hyper-parameters depended heavily on the choice of metric used during the calibration phase, where data assimilation was applied. These sets of hyper-parameters in turn produced different hydrologic forecasts. This influence was reduced as the forecast lead time increased, because of not applying data assimilation in the forecast mode, and accordingly, convergence of model state ensembles produced in the calibration phase. However, the influence could remain considerable for a few days up to multiple weeks depending on the catchment and the model. As such, a preliminary analysis would be recommended for future studies to better understand the impact that metrics can have within and outside the bounds of hyper-parameter calibration.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.061 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".