On the exploration of alternative spatial representation for land models; a vector-based setup for the Variable Infiltration Capacity model
Bibliographic record
Abstract
Land models are increasingly used as the backbone of the terrestrial hydrology as they cover a wide range of processes (from rainfall/runoff processes to carbon cycle). The recent improvements in high-resolution spatial data set including detailed digital elevation models, DEMs, and land cover and soil type maps are encouraging the modelers to set up the land surface models at the highest resolution possible. However, this high-resolution setup does not often coincide with rigorous model diagnostics and also the “optimal” spatial representation based on the context of modeling (e.g. streamflow). A model can be seen as a tool to interpolate or extrapolate our knowledge in time and space and therefore it remains an important aspect of land surface modeling to which level the spatial heterogeneity can be represented in a model so that the states and fluxes “improve” given the context of modeling. The representation of spatial data in our models has important implications including (1) removing the unnecessarily computational burden from model setups which in turn results in better assessment of uncertainty and sensitivity analysis of the parameters on a less computational expensive model. (2) Proper corresponding between the communications of spatial variability while avoiding overconfidence in the nature of model response on illogically smallest units. In this study, in contrast to the often used grid-based model setup, we use the concept of vector-based group response units (GRUs) for setting up the Variable Infiltration Capacity, the VIC model, and vector-based MizuRoute routing scheme. We explore the added information by stepwise inclusion of more detailed spatial data and higher resolution forcing data while the vector-based routing setup remains identical for each of the configurations. Using this flexible workflow we explore three major questions: 1- How the performance of model changes in the calibration mode for various configuration of spatial heterogeneity representation and forcing resolution given the context of modeling, for example, streamflow simulations or snow water equivalent spatial pattern? 2- How well a simplified version of a more complex model in spatial representation can reproduce its own simulation? The answer to this question will provide us with iso-performing model setups, configurations of forcing distribution and spatial heterogeneity representation, and the possible loss in the performance metric given the context of modeling under the simplification decisions. 3- How the model performs across various configurations of spatial data and forcing resolutions with a given set of so-called physically parameters that are often considered to be identical for GRUs with the same physical characteristics, soil, vegetation type, elevation zone, slope and aspect, varies? Our findings indicate that the optimal spatial representation in the context of modeling, streamflow, for example, may very well be much less computationally demanding than the model setup that contains all the details with the highest resolution of the data. In a complementary attempt, it is shown that the often good performing parameter sets are able to reproduce good performing simulation in comparison to the model setup with the highest model resolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".