What is the role of machine learning when we want to simulate hydrological processes?
Bibliographic record
Abstract
It has now been almost five years since Grey Nearing and his colleagues published their provocative commentary “What Role Does Hydrological Science Play in the Age of Machine Learning?”. Nearing et al. reviewed experiments that use deep learning to simulate time series of streamflow, emphasizing results that show there is substantially more information in large‐domain hydrological data sets than hydrologists have been able to translate into theory or models. In their commentary, Nearing et al. encouraged the hydrology community “to focus on developing a quantitative understanding of where and when hydrological process understanding is valuable in a modeling discipline [that is] increasingly dominated by machine learning.”This presentation will summarize advances in process-based hydrological modeling in our research group in the five years since Nearing et al. published their controversial commentary. To bridge the gap between process-based modeling and machine learning, we depart from the focus of Nearing et al. where machine learning has a central role in the modeling ecosystem – instead, we ask how machine learning can enable and accelerate the development of process-based hydrological models. We will emphasize the components of the model ecosystem where we use machine learning and artificial intelligence, and the ecosystem components where we do not. We will discuss our advances in generating ensemble spatial meteorological fields, the numerical implementation of process-based models, process-based parameter estimation, multi-model combinations, and reproducible and transparent workflows. We will demonstrate tangible progress in closing the gap between the predictive performance of (hybrid) process-based models and pure machine learning algorithms for hydrological predictions across large geographical domains. We also demonstrate prototype workflows that use artificial intelligence to support the hydrological modelling exercise from A-Z, including the configuration, running, optimisation and interpretation of complex process-based models. We consider the community value and dangers of using AI to assist in different aspects of the process of scientific discovery.We will end the presentation by returning to the question posed by Nearing et al. – What Role Does Hydrological Science Play in the Age of Machine Learning? We will argue that the appropriate use of machine learning and artificial intelligence is beginning to enable the development of process-based models that effectively use the information in large-domain hydrological datasets, while maintaining the interpretability and transparency of physically grounded simulations. We will suggest a path forward for the discipline where machine learning and artificial intelligence are essential to develop the next generation of hydrological prediction systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.105 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.019 |
| Scholarly communication | 0.008 | 0.019 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.016 | 0.033 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".