Evaluating global hydrological-process modelling beyond river discharge observations
Bibliographic record
Abstract
Catchment modelling of water balance components is nowadays done at high spatial resolution for continental and global scales, thanks to the increasing computational capacity and the growing trend towards open data. One of these process-based models is the World-Wide HYPE (WW-HYPE; Arheimer et al., 2020), which was set-up by a stepwise calibration strategy to avoid equifinality when using streamflow data for parameter estimation. In this presentation we suggest to further evaluate whether the model is right for the right reason by comparing internal variables against independent Earth Observations (EO). We then assume that the results are robust if the two different sources of data reveal the same results. This approach could become a new standard method today for evaluating continuous process-based global models as there are numerous EO products representing various hydrological variables, most of them covering at least the last decade.We propose to compare three aspects when evaluating robustness in global hydrological variables: i) long-term means, ii) seasonal variability through monthly means, and iii) equifinality by comparing model-streamflow performance versus internal variable performance.We applied this method by comparing six hydrological variables (potential and actual evapotranspiration, snow cover, snow water equivalent, soil moisture or changes in water storages) from EO-products (based on MODIS, GlobSnow, ESA-CCI Soil Moisture and GRACE) with WWH variables for the time-period 2000-2014 (Pimentel et al, 2023). We then found that the general patterns in the hydrological cycle show good agreement between catchment modelling and EO at the global scale, although some months in water-storage changes differed. These dissimilarities indicate that hydrological variables above the ground and earlier in the flow path are more robust than the sub-surface downstream processes, such as soil moisture distribution and water-storage changes, which reflect more complex processes that can be challenging to describe both by hydrological models and satellite sensors. Regarding geographical distribution, there is a larger spread in results from regions with extreme characteristics, such as cold regions (Canadian prairies), arid regions (western USA, deserts), highly forested areas (Amazonas), and transition zones (Sahel and Mediterranean Basin). This indicate that the particularity of these regions calls for specific regional modelling and monitoring approaches rather than continental or global approaches.On the contrary, in temperate regions at mid-latitudes, e.g., eastern USA and central Europe, almost all the hydrological variables were found robust. With respect to equifinality, overall, there were no indication on good discharge performance and bad internal model representation. The exercise shows the potential in using EO products for model evaluation beyond traditional river-discharge observations from gauges, to first assess the robustness of hydrological variables and second to determine which processes should be better represented in model parameterisation, without forgetting that EO products are not a ground truth and are also assigned with uncertainties. References:Arheimer et al., 2020: Global catchment modelling using World-Wide HYPE (WWH), open data and stepwise parameter estimation, HESS 24, 535–559, https://doi.org/10.5194/hess-24-535-2020Pimentel et al., 2023: Assessing Robustness in Global Hydrological Modelling through EO Comparisons, HSJ (in review)
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".