On doing large-scale hydrology with Lions: Realising the value of perceptual models and knowledge accumulation
Bibliographic record
Abstract
Moving the study domain in hydrology to larger and larger regions leaves us with significant knowledge gaps because we are unable to observe the hydrology of many parts of the world, while in-depth hydrologic studies cover only a fraction of our landscape. On medieval maps, knowledge gaps were shown as images of lions. How do we best acknowledge and reduce these gaps in hydrology, i.e. our hydrologic lions? The accumulation of knowledge has been postulated as the fundamental mark of scientific advancement by some philosophers of science. In hydrology, knowledge accumulation has been somewhat fragmented, left as a pursuit for (often brilliant) individuals rather than emphasised as a necessary focus for the research community. Our knowledge of a region’s hydrology originates from available observations. However, the ability of observations to reliably characterise hydrological phenomena is limited, and large areas of the globe lack detailed observations. In this commentary we propose two strategies to rectify these deficiencies. First, the use of shared perceptual models as ways to capture, debate and test our experience with different hydrologic systems. Second, improved knowledge accumulation in hydrology by more strongly focusing on knowledge extraction from available historical articles. This effort should include the addition of meta-data to tag hydrologic journal articles and by developing a related hydrological database that would enable searching, organizing and analysing previous studies in a hydrologically meaningful manner.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.041 | 0.087 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.004 | 0.029 |
| Scholarly communication | 0.019 | 0.041 |
| Open science | 0.004 | 0.010 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".