Global prevalence of non-perennial rivers and streams
Bibliographic record
Abstract
<b>Global prevalence of non-perennial rivers and streams</b>June 2021<b><br></b>prepared by <b>Mathis L. Messager (mathis.messager@mail.mcgill.ca)</b><b>Bernhard Lehner (bernhard.lehner@mcgill.ca)</b><br>1. Overview and background 2. Repository content3. Data format and projection4. License and citations4.1 License agreement4.2 Citations and acknowledgements<br><br><b>1. Overview and background</b>This documentation describes the data produced for the research article: Messager, M. L., Lehner, B., Cockburn, C., Lamouroux, N., Pella, H., Snelder, T., Tockner, K., Trautmann, T., Watt, C. & Datry, T. (2021). Global prevalence of non-perennial rivers and streams. Nature. https://doi.org/10.1038/s41586-021-03565-5<br>In this study, we developed a statistical Random Forest model to produce the first reach-scale estimate of the global distribution of non-perennial rivers and streams. For this purpose, we linked quality-checked observed streamflow data from 5,615 gauging stations (on 4,428 perennial and 1,187 non-perennial reaches) with 113 candidate environmental predictors available globally. Predictors included variables describing climate, physiography, land cover, soil, geology, and groundwater as well as estimates of long-term naturalised (i.e., without anthropogenic water use in the form of abstractions or impoundments) mean monthly and mean annual flow (MAF), derived from a global hydrological model (WaterGAP 2.2; Müller Schmied et al. 2014). Following model training and validation, we predicted the probability of flow intermittence for all river reaches in the RiverATLAS database (Linke et al. 2019), a digital representation of the global river network at high spatial resolution.<br>The data repository includes two datasets resulting from this study:1. a geometric network of the global river system where each river segment is associated with:i. 113 hydro-environmental predictors used in model development and predictions, andii. the probability and class of flow intermittence predicted by the model.2. point locations of the 5,516 gauging stations used in model training/testing, where each station is associated with a line segment representing a reach in the river network, and a set of metadata.<br>These datasets have been generated with source code located at messamat.github.io/globalirmap/.<br>Note that, although several attributes initially included in RiverATLAS version 1.0 have been updated for this study, the dataset provided here is not an established new version of RiverATLAS. <br><br><br><b>2. Repository content</b>The data repository has the following structure (for usage, see section 3. Data Format and Projection; GIRES stands for Global Intermittent Rivers and Ephemeral Streams):<br>— <i><b>GIRES_v10_gdb.zip/ </b>: file geodatabase in ESRI® geodatabase format containing two feature classes (zipped)</i> |——— <b>GIRES_v10_rivers</b> : river network lines |——— <b>GIRES_v10_stations</b> : points with streamflow summary statistics and metadata<br>—<b> </b><i><b>GIRES_v10_shp.zip/ </b>: directory containing ten shapefiles (zipped)</i> Same content as GIRES_v10_gdb.zip for users that cannot read ESRI geodatabases (tiled by region due to size limitations). |——— <b>GIRES_v10_rivers_af.shp</b> : Africa |——— <b>GIRES_v10_rivers_ar.shp</b> : North American Arctic |——— <b>GIRES_v10_rivers_as.shp</b> : Asia |——— <b>GIRES_v10_rivers_au.shp</b> : Australasia|——— <b>GIRES_v10_rivers_eu.shp</b> : Europe|——— <b>GIRES_v10_rivers_gr.shp</b> : Greenland|——— <b>GIRES_v10_rivers_na.shp </b>: North America|——— <b>GIRES_v10_rivers_sa.shp</b> : South America<br>|——— <b>GIRES_v10_rivers_si.shp</b> : Siberia<br>|——— <b>GIRES_v10_stations.shp</b> : points with streamflow summary statistics and metadata<br>— <b><i>Other_technical_documentations.zip/</i></b> :<i> directory containing three documentation files (zipped)</i>|——— <b>HydroATLAS_TechDoc_v10.pdf</b> : documentation for river network framework|——— <b>RiverATLAS_Catalog_v10.pdf</b> : documentation for river network hydro-environmental attributes|——— <b>Readme_GSIM_part1.txt</b> : documentation for gauging stations from the Global Streamflow Indices and Metadata (GSIM) archive<br>—<b> README_Technical_documentation_GIRES_v10.pdf </b>: full documentation for this repository<b><br></b><b><br></b><b>3. Data format and projection</b>The geometric network (lines) and gauging stations (points) datasets are distributed both in ESRI® file geodatabase and shapefile formats. The file geodatabase contains all data and is the prime, recommended format. Shapefiles are provided as a copy for users that cannot read the geodatabase. Each shapefile consists of five main files (.dbf, .sbn, .sbx, .shp, .shx), and projection information is provided in an ASCII text file (.prj). The attribute table can be accessed as a stand-alone file in dBASE format (.dbf) which is included in the Shapefile format. <br>These datasets are available electronically in compressed zip file format. To use the data files, the zip files must first be decompressed.<br>All data layers are provided in geographic (latitude/longitude) projection, referenced to datum WGS84. In ESRI® software this projection is defined by the geographic coordinate system GCS_WGS_1984 and datum D_WGS_1984 (EPSG: 4326).<br><b><br></b><b>4. License and citations</b><i>4.1 License agreement </i>This documentation and datasets are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License (CC-BY-4.0 License). For all regulations regarding license grants, copyright, redistribution restrictions, required attributions, disclaimer of warranty, indemnification, liability, waiver of damages, and a precise definition of licensed materials, please refer to the License Agreement (https://creativecommons.org/licenses/by/4.0/legalcode). For a human-readable summary of the license, please see https://creativecommons.org/licenses/by/4.0/.<br><br><i>4.2 Citations and acknowledgements.</i>Citations and acknowledgements of this dataset should be made as follows:Messager, M. L., Lehner, B., Cockburn, C., Lamouroux, N., Pella, H., Snelder, T., Tockner, K., Trautmann, T., Watt, C. & Datry, T. (2021). Global prevalence of non-perennial rivers and streams. Nature. https://doi.org/10.1038/s41586-021-03565-5 <br>We kindly ask users to cite this study in any published material produced using it. If possible, online links to this repository (https://doi.org/10.6084/m9.figshare.14633022) should also be provided.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.441 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".