Classifying annual daily hydrographs in Western North America using <scp>t‐distributed</scp> stochastic neighbour embedding
Bibliographic record
Abstract
Abstract Flow regimes are critical for determining physical and biological processes in rivers and their classification and regionalization traditionally seeks to link patterns of flow to physiographic, climate and other information. There are many approaches to, and rationales for, catchment classification, with those focused on streamflow often seeking to relate a particular response characteristic to a physical property or climatic driver. Rationales include such topics as prediction in ungauged basins (PUB), and providing guidance for model selection in poorly understood hydrological systems. The annual daily hydrograph (ADH) is a first‐order easily visualized integrated expression of catchment function, and over many years the average ADH is a distinct hydrological signature that differentiate catchments from each other. In this study, we use t‐SNE, a state‐of‐the‐art technique of dimensionality reduction, to classify 17 110 ADHs for 304 reference catchments in mountainous Western North America. t‐SNE is chosen over other conventional methods of dimensionality reduction (e.g., PCA) as it presents greater separability of ADHs, which are projected on a 2D map where the similarities are evaluated according to their map distance. We then utilize a Deep Learning encoder to upgrade the non‐parametric t‐SNE to a parametric approach, enhancing its capability to address ‘unseen’ samples. Results showed that t‐SNE successfully clustered ADHs of similar flow regimes on the 2D map and allowed more accurate classification with KNN. In addition, many compact clusters on the 2D map in the coastal Pacific Northwest suggest information redundancy in the local streamflow network. The t‐SNE map provides an intuitive way to visualize the similarity of high‐dimensional data of ADHs, groups catchments with like characteristics, and avoids the reliance on subjective hydrometric indicators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".