SWOT River Database (SWORD)
Bibliographic record
Abstract
SUPPLEMENTAL FILE NOTE (2026-09-02): Added SWORD_v17b_to_v16_Translation.zip, containing the official global reach and node ID translation tables for all six regions. The underlying SWORD v17c data products are unchanged. VERSION NOTES: v17c versus v17b SWORD v17c is an enhancement release built on v17b by the Global Hydrology Lab at the University of North Carolina at Chapel Hill. SWORD v17b remains the official basis for SWOT Version D RiverSP Vector Products; v17c is not the JPL-official product version. v17c preserves the full v17b structure — 248,673 reaches, 11,112,454 nodes, and ~66.9 million centerline points across 6 continental regions — with reach and node counts identical to v17b. Reach topology (neighbor relationships) and reach distance-from-outlet (dist_out) are likewise identical to v17b. v17c retains every v17b variable (no columns removed) and adds new columns on top. The v17b static fields are unchanged; the only changes to existing variables are the facc correction, the node dist_out midpoint convention, and the lakeflag/type reconciliation described below. The main additions: Actual SWOT-derived observational data (new): for the first time, SWORD carries measured SWOT water-surface elevation, width, and slope aggregated per reach and node from real overpasses: percentile summaries (wse_obs_*, width_obs_*, slope_obs_*), spread statistics, and observation counts (n_obs). These are new columns; the v17b static wse, width, and slope variables are retained unchanged. Flow-accumulation correction: facc was corrected on 95,880 reaches (38.6%) to remove three systematic MERIT Hydro D8 routing artifacts — bifurcation cloning, junction double-counting, and raster–vector misalignment — raising 80,538 reaches by a median of 60% and lowering 15,342 by a median of 30% relative to v17b, and eliminating all junction-conservation and downstream-monotonicity violations. A per-reach flag (facc_quality) marks the corrected reaches. Computed topology and routing (new): mainstem identification (is_mainstem, main_path_id, rch_id_up_main/rch_id_dn_main), best-headwater/outlet routing (best_headwater, best_outlet), new distance metrics (dist_out_dijkstra, hydro_dist_out, pathlen_hw/pathlen_out), and a globally unique connected-component identifier (subnetwork_id). Data quality fixes to v17b-inherited fields: node geolocation repairs, node ordering normalization, lakeflag/type reconciliation (including HarP-informed lake corrections), and lake-sandwich corrections. Node dist_out now uses a midpoint convention (each node's distance refers to the center of the ~200 m segment it represents), consistent with the new v17c distance variables. Formats: v17c is distributed in NetCDF, GeoPackage, Shapefile, GeoParquet, and DuckDB. Files are packaged one ZIP per format, split by continental region (Shapefiles are further split by HydroBASINS level-2 basin). Whole-planet merged GeoParquet and DuckDB files are also provided in the "global" ZIP. Please cite this dataset. When you use SWORD v17c in your work, cite this Zenodo record: Gearon, J. H., et al. (2026). SWOT River Database (SWORD), version v17c [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3898569 Alternatively, you may cite the original SWORD paper: Altenau, E. H., Pavelsky, T. M., Durand, M. T., Yang, X., Frasson, R. P. d. M., & Bendezu, L. (2021). The Surface Water and Ocean Topography (SWOT) Mission River Database (SWORD): A global river network for satellite data products. Water Resources Research, 57(7), e2021WR030054. https://doi.org/10.1029/2021WR030054 The DOI 10.5281/zenodo.3898569 always resolves to the most recent version of SWORD; a version-specific DOI for v17c is shown on this record's page for citing this exact release. 1. Summary: The Surface Water and Ocean Topography (SWOT) satellite mission vastly expands observations of river water surface elevation (WSE), width, and slope. The SWOT River Database (SWORD) combines multiple global river- and satellite-related datasets to define the nodes and reaches that constitute SWOT river vector data products. SWORD provides high-resolution river nodes (200 m) and reaches (~10 km) with attached hydrologic variables (WSE, width, slope, etc.) as well as a consistent topological system for global rivers 30 m wide and greater. 2. Data Formats: SWORD v17c is provided in netCDF, GeoPackage, shapefile, GeoParquet, and DuckDB formats. All files use a two-letter continent identifier ("af" – Africa, "as" – Asia / Siberia, "eu" – Europe / Middle East, "na" – North America, "oc" – Oceania, "sa" – South America). NetCDF files are structured in 3 groups (centerlines, nodes, and reaches) and are distributed at continental scales. GeoPackage, GeoParquet, and DuckDB files carry nodes and reaches per continental region, where nodes are ~200 m spaced points and reaches are polylines. All vector files are in geographic (latitude/longitude) projection, referenced to datum WGS84. 3. Attribute Description: The catalog below documents every variable in the SWORD v17c files, grouped into the new v17c variables and the inherited v17b variables, for the reaches, nodes, and centerlines groups. Fill values: i4 = -9999, i8 = -9999, f8 = -9999.0. The reaches edit_flag variable declares a string _FillValue attribute of "-9999.0", but reaches with no edit tag are stored as empty strings (""), not the literal "-9999.0"; other string variables have no fill value. In the rch_id_up / rch_id_dn [4, N] neighbor arrays, empty slots use -9999 (v17b used 0). Distance Variable Conventions SWORD contains four distance-to-outlet variables. They differ in routing method, zero-point, and reach endpoint reporting convention. The table below summarizes these conventions; see "Node-Level Interpolation" for how each extends to nodes. Variable Routing Reach-level value Outlet value Ghost reaches dist_out (v17b) v17b topology reported at reach upstream endpoint reach_length has value hydro_dist_out rch_id_dn_main chain reported at reach upstream endpoint reach_length has value dist_out_dijkstra Dijkstra shortest path reported at reach upstream endpoint reach_length has value hydro_dist_hw rch_id_up_main chain reported at reach upstream endpoint 0 has value Reach scalar convention. These reach-level values are endpoint reports, not traversal-origin claims. dist_out, hydro_dist_out, and dist_out_dijkstra are outlet-distance scalars reported at the upstream endpoint of each reach; each reach exceeds its downstream neighbor by its own reach_length on 1:1 links. hydro_dist_hw is also reported at the upstream endpoint, but measures distance from best_headwater; each downstream reach exceeds its upstream neighbor by the upstream neighbor's reach_length on 1:1 links. Reference values. dist_out, hydro_dist_out, and dist_out_dijkstra all assign an outlet reach a value equal to its reach_length (the upstream endpoint is one reach length from the outlet point). hydro_dist_hw assigns 0 at the headwater. Ghost reaches. All four distance variables retain values for ghost reaches in v17c, including dist_out_dijkstra. Node-Level Interpolation Six distance variables are interpolated per node using a midpoint offset within the parent reach: offset = cumsum(node_length) - 0.5 * node_length, where the cumulative sum is ordered by node_order. This places each node at the geometric center of its node_length segment. On a single-path network (no junctions), dist_out, hydro_dist_out, and dist_out_dijkstra are exactly equal at every node. Variable Node formula node_order=n (upstream) node_order=1 (downstream) dist_out reach.do - reach_length + offset ~ reach.do ~ reach.do - reach_length hydro_dist_out reach.hdo - reach_length + offset ~ reach.hdo ~ reach.hdo - reach_length dist_out_dijkstra reach.dod - reach_length + offset ~ reach.dod ~ reach.dod - reach_length hydro_dist_hw reach.hdh + reach_length - offset ~ reach.hdh ~ reach.hdh + reach_length pathlen_hw reach.plh - reach_length + offset ~ reach.plh - reach_length ~ reach.plh pathlen_out reach.plo + reach_length - offset ~ reach.plo + reach_length ~ reach.plo Three additional variables (subnetwork_id, best_headwater, best_outlet) are flat copies from the parent reach. Boundary behavior. The reach-level formulas imply continuous endpoint values on 1:1 links, but node values are midpoint samples within their node_length segments. Adjacent boundary nodes therefore need not be identical, and bifurcation-rejoin structures can retain larger single-scalar dist_out gaps. New v17c Reach Variables Variable NetCDF Type Units Fill Value Encoding Description dist_out_dijkstra f8 meters -9999.0 Dijkstra shortest-path outlet distance reported at the reach upstream endpoint; outlet reach = reach_length; values retained for ghost reaches hydro_dist_out f8 meters -9999.0 Mainstem outlet distance reported at the reach upstream endpoint via rch_id_dn_main chain; outlet reach = reach_length hydro_dist_hw f8 meters -9999.0 Mainstem headwater distance reported at the reach upstream endpoint via rch_id_up_main chain; headwater = 0 rch_id_up_main i8 -9999 Main upstream neighbor reach ID (mainstem-preferred) rch_id_dn_main i8 -9999 Main downstream neighbor reach ID (mainstem-preferred) subnetwork_id i4 -9999 Connected component ID (Pfafstetter-offset, globally unique; differs from v17b network) main_path_id i8 -9999 ID of the mainstem path this reach belongs to is_mainstem i4 -9999 0=not_mainstem; 1=mainstem Whether reach is on a mainstem path best_headwater i8 -9999 Width-prioritized upstream headwater reach ID best_outlet i8 -9999 Width-prioritized downstream outlet reach ID pathlen_hw f8 meters -9999.0 Cumulative reach_length sum from best_headwater to the downstream end of the reach (headwater = 0, increases downstream) pathlen_out f8 meters -9999.0 Cumulative reach_length sum toward best_outlet (outlet = 0, increas
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.204 | 0.314 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".