On the elaboration of a robust calibration strategy for the large-scale GEM-Hydro model
Bibliographic record
Abstract
As part of the Great-Lakes Runoff Inter-comparison Project (GRIP-GL; Mai et al., 2022), which aims at comparing the performances of different hydrologic models over the Great-Lakes when calibrating them using the same meteorological inputs and geophysical databases, the GEM-Hydro hydrologic model used at Environment and Climate Change Canada (ECCC) to perform operational hydrologic forecasts was calibrated using different strategies. Following the calibration work related to GRIP-GL, progress has been achieved with regard to improving the calibration of the GEM-Hydro model.The work presented here focuses on improvements achieved with regard to calibrating the GEM-Hydro model, compared to the default version of the model and to the performances obtained during the GRIP-GL project. For various reasons explained, the GEM-Hydro calibration performed as part of GRIP-GL was suboptimal. The general calibration framework remains the same as in GRIP-GL, for example by using the MESH-SVS-Raven model to speed-up simulation times and transferring the calibrated parameters into GEM-Hydro afterwards, by relying on global calibrations for each of the 6 Great-Lakes subdomains, etc. However, several important changes have been made compared to the work performed in GRIP-GL, like a new approach to represent the effect of Tile Drains, changing the set of flow stations used for calibration, revising the objective function, etc.The proposed calibration methodology updates significantly improve GEM-Hydro streamflow performance across the Great-Lakes domain and in addition also improve or maintain similar performance levels as the default version of the model, with respect to auxiliary variables and surface fluxes: snow, soil moisture, evapotranspiration, 2m air temperature and dew point. Indeed, the model relies on 40m atmospheric forcings for wind speed, temperature and humidity, and simulates its own 2m atmospheric variables. To achieve this, it was necessary to constrain some parameter interval values during calibration, in order to prevent the calibration algorithm to choose physically-irrelevant parameter values that could allow to improve streamflow performances while degrading other hydrologic variables, due to equifinality.Reference:Mai, J., Shen, H., Tolson, B. A., Gaborit, E., Arsenault, R., Craig, J. R., Fortin, V., Fry, L. M., Gauch, M., Klotz, D., Kratzert, F., O'Brien, N., Princz, D. G., Rasiya Koya, S., Roy, T., Seglenieks, F., Shrestha, N. K., Temgoua, A. G. T., Vionnet, V., and Waddell, J. W. (2022). The Great Lakes Runoff Intercomparison Project Phase 4: The Great Lakes (GRIP-GL). Hydrol. Earth Syst. Sci., 26, 3537–3572. Highlight paper. Accepted Jun 10, 2022. https://doi.org/10.5194/hess-26-3537-2022
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".