Empirical Evidence of the Importance of Data Recency in LSTM-Based Rainfall-Runoff Modeling 
Bibliographic record
Abstract
Deep learning (DL)-based hydrological models, particularly those using Long Short-Term Memory (LSTM) networks, typically require large datasets for effective training. In the context of large-scale rainfall-runoff modeling, dataset size can refer to either the number of watersheds or the length of the training period. While it is well established that training a regional model across more watersheds improves performance (Kratzert et al., 2024), the benefits of extending the training period are less clear.Empirical evidence from studies such as Boulmaiz et al. (2020) and Gauch et al. (2021) suggests that longer training periods enhance LSTM performance in rainfall-runoff modeling. This improvement is attributed to the need for extensive datasets to ensure proper model convergence and the ability to capture a wide range of hydrological conditions and events. However, these studies neglected the influence of data recency (or data recentness), which is critical for operational applications that forecast current and future hydrological conditions. In the context of climate change and anthropogenic interventions, the assumption of stationarity (i.e., that historical patterns reliably represent future conditions) may no longer hold for hydrological systems (Shen et al., 2022). Consequently, the selection of training periods should account for potential non-stationarity, as more recent data may better reflect current rainfall-runoff dynamics. Intriguingly, Shen et al. (2022) found that calibrating hydrologic models to the latest data is a superior approach compared to using old data, and completely discarding the oldest data can even improve the performance in streamflow prediction.This study aims to address two research questions: (1) As the number of watersheds increases, is it still necessary to train LSTM models on decades of historical observations? (2) Can LSTM models achieve comparable performance using shorter training periods focused on more recent data? Specifically, we examine whether models trained on recent data outperform those trained on older data and explore how different temporal partitions of historical records affect predictive skill.This study leverages a comprehensive dataset comprising streamflow records from over 1,300 watersheds across North America, representing diverse climatic and hydrological regimes, with streamflow data spanning 1950 to 2023. Training periods are designed to isolate the effects of temporal data recency while keeping period lengths consistent. This approach enables a systematic comparison of model performance using exclusively older (e.g., pre-1980) versus exclusively recent data (e.g., post-1980). This research provides evidence-based recommendations for selecting training data while balancing computational costs, data availability, and prediction accuracy. ReferencesBoulmaiz, T., Guermoui, M., and Boutaghane, H.: Impact of training data size on the LSTM performances for rainfall–runoff modeling, Model Earth Syst Environ, 6, 2153–2164, https://doi.org/10.1007/S40808-020-00830-W/FIGURES/9, 2020.Gauch, M., Mai, J., and Lin, J.: The proper care and feeding of CAMELS: How limited training data affects streamflow prediction, Environmental Modelling and Software, 135, https://doi.org/10.1016/j.envsoft.2020.104926, 2021.Kratzert, F., Gauch, M., Klotz, D., and Nearing, G.: HESS Opinions: Never train an LSTM on a single basin, Hydrology and Earth System Science, https://doi.org/10.5194/hess-2023-275, 2024.Shen, H., Tolson, B. A., and Mai, J.: Time to Update the Split-Sample Approach in Hydrological Model Calibration, Water Resour Res, 58, e2021WR031523, https://doi.org/10.1029/2021WR031523, 2022.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.051 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.002 | 0.005 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".