Canada1Water classification of the National Hydro Network: stream order and graph refinement
Bibliographic record
Abstract
A vector representation of stream networks is a crucial dataset for the modelling the surface water and groundwater components of the hydrologic cycle. For many usages a crucial attribute of the drainage network is a digital topology and hierarchal stream order attribute (e.g., Strahler stream order). In Canada jurisdictional stream networks are available for the provinces and territories and nationally for Canada in the National Hydrological Network (NHN) dataset. Unfortunately, the NHN data lacks the same topological and attribute information that is available for numerous provinces due to standardization for the entire country. For Canada1Water it was also necessary to have a harmonized dataset with the United States, for both the southern transboundary watersheds and the Alaskan watersheds. This report documents the processes completed to upgrade the topological and graph network support for NHN and provide continuous connectivity with US datasets. It also highlights and corrects a number of stream density and stream order issues that occur within Canada across provincial and territorial borders and NTS tiles. All vector processing was completed in RivEX software extension for ArcMap. Following complete topological correction stream classification was assigned and a table of the node graph network developed. Additional work was then completed to normalize stream density particularly amongst low-order streams between British Columbia and the Yukon and amongst local NTS tiles in Quebec and Ontario. Corrected NHN Strahler stream order assignment was validated against a number of provincial and watershed datasets, all of which already have Strahler stream order attributed. These datasets are the same underlying digitized vector data, so there are no differences in node or polyline positions. Strahler stream order assignment validation was only done by visual comparison as due to differences in vector segments a statistical comparison is complicated. The transboundary integrated C1W stream network with complete classification provides a seamless national dataset to support transdisciplinary studies (fisheries, wildlife, health, pesticide and nutrient issues, mining impact, ecosystem restoration, numeric modelling) that involve a knowledge of stream distribution and ranking.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.011 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.008 | 0.015 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.005 | 0.003 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.099 | 0.030 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".