Time to Update the Split Sample Approach to Hydrological Model Calibration: A Massive Empirical Study
Bibliographic record
Abstract
Model calibration and validation are critical in hydrological model robustness assessment. Unfortunately, the commonly used split-sample test (SST) framework for data splitting requires modelers to make subjective decisions without clear guidelines. A massive SST experiment for hydrological modeling is proposed and tested across a large sample of catchments to empirically reveal how data availability and calibration period features (i.e., length and recentness) simultaneously impact model performance in the post-validation period (e.g., forecasting or prediction), thus providing practical guidance on split-sample design. Unlike most SST studies that use two sub-periods (i.e., calibration and validation) to build models, this study incorporates an independent model testing period in addition to calibration and validation periods. Model performance of two lumped conceptual hydrological models (i.e., GR4J and HMETS) are calibrated and tested in 463 CAMELS catchments across the United States using 50 different data splitting schemes. These schemes are established regarding the data availability, length, and data recentness of the continuous calibration sub-periods (CSPs). A full-period CSP is also included in the experiment, which skips model validation entirely. The results are synthesized regarding the large sample of catchments and are comparatively assessed in multiple novel ways, including how model building decisions are framed as a decision tree problem and viewing the model validation process as a formal testing period classification problem, aiming to accurately predict model success/failure in the testing period. Results span different climate and catchment conditions across a 35-year period with available data, making conclusions generalizable. Strong patterns show that calibrating to older data and then validating models on newer data produces inferior model testing period performance in every single analysis conducted and should hence be avoided. Calibrating to the full available data and skipping model validation entirely is the most robust split-sample decision. Findings have significant implications for SST practice in hydrological modeling. As the next phase of this study, results for discontinuous calibration sub-periods (DCSP) will be evaluated as an alternative SST design choice and contrasted then with the CSP results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.031 | 0.161 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.005 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".