Bibliographic record
Abstract
<b>HydroFATE (v1): A high-resolution contaminant fate model for the global river system</b>Last updated: December 2023prepared byHeloisa Ehalt Macedo (heloisa.ehaltmacedo@mail.mcgill.ca) and Bernhard Lehner (bernhard.lehner@mcgill.ca)<b>1. Overview and background</b>This documentation describes the input data necessary for the operation of the global contaminant fate model HydroFATE, available at https://doi.org/10.6084/m9.figshare.23646282. In addition, it describes the results from a case study presented at the research article: Ehalt Macedo, H., Lehner, B., Nicell, J., Grill, G. (2024). HydroFATE (v1): A high-resolution contaminant fate model for the global river system. Further information and description of the model can be found at the same publication.The data repository includes 4 datasets:1. Python code: python project repository including the structure necessary for the model to run.2. Input data:a. A table containing information on all river segments associated with geometric attributes from RiverATLAS (Linke et al., 2019) and HydroROUT (Lehner and Grill, 2013), and attributes used in the HydroFATE model based on underlying data such as HydroLAKES (Messager et al., 2016), and HydroWASTE (Ehalt Macedo et al., 2022).b. A table of parameters related to the contaminant that is being analyzed and their configuration for the model run (scenarios). For the case study, the literature sources of the parameters for SMX and the scenarios are described in the research paper.c. A table of country-level consumption per capita of the contaminant that is being analyzed. For the case study, the consumption information for Sulfamethoxazole (SMX) was provided by Klein et al. (2018).3. Output data:a. A table including the unique river reach identifier associated with the river network and the resulting concentration of SMX using HydroFATE for the 4 main scenarios described at the research paper.4. Contaminant pathways dataset:a. A grid containing a unique identifier associated with wastewater treatment plants (WWTPs) from the database HydroWASTE, decentralized wastewater systems (DWTS), urban untreated, or rural untreated; i.e., the different contaminant pathways.b. A table including the unique identifier associated with the WWTP and additional attributes, such as population assigned.c. Python code: python project repository including the structure necessary for the tool created to delineate wastewater treatment plants (WWTP) service areas in order to define contaminant pathways for the global contaminant fate model HydroFATE.<br><b>2. Repository content</b>The data repository has the following structure:<b>HydroFATE_v1.zip/: </b>repository containing:<b>|---------Main_script/:</b><b>|---------------------Input/</b>: empty folder to add input data from “Input_data.gdb.zip”<b>|---------------------</b><b>Output/</b>: empty folder where results will be saved after model run<b>|---------------------</b><b>HydroFATE_v1.py</b> : python code with HydroFATE model<b>|---------------------</b><b>config.py:</b> config file with model parameters<b>|---------</b><b>LICENSE</b>: license file for python code<b>|---------</b><b>README.md</b>: readme file for code description and compilation instructions<b>Input_data.gdb.zip/:</b> file geodatabase in ESRI® geodatabase format containing 3 feature classes (zipped):<b>|---------</b><b>streams:</b> table including global river network attributes.<b>|---------</b><b>parameters</b>: table including parameters and configuration settings.<b>|---------</b><b>consumption: </b>table including country-level consumption per capita of substance.<b>Output.gdb.zip/:</b> file geodatabase in ESRI® geodatabase format containing 1 feature classes (zipped):<b>|---------</b><b>SMX_results</b>: table with model predictions for the case study for every river reach of the global river network.<b>Contaminant_Pathways.zip/: </b>repository containing:<b>|---------</b><b>cont_pathways.tif:</b> grid in TIFF format<b>|---------</b><b>HydroWASTE_SA.csv: </b>table in CSV format<b>WWTP_delineation.zip/: </b>repository containing:<b>|---------</b><b>Main_script/:</b><b>|---------------------</b><b>Input/</b>: empty folder to add input data to run the tool<b>|---------------------</b><b>Output/</b>: empty folder where results will be saved after model run<b>|---------------------</b><b>WTP_delineation.py</b> : python code with tool to delineate WWTP service areas<b>|---------</b><b>LICENSE</b>: license file for python code<b>|---------</b><b>README.md</b>: readme file for code description and compilation instructions<b>3. Data format and projection</b>A license for the software ArcGIS Pro is required to run the provided scripts. These datasets are available electronically in compressed zip file format. To use the data files, the zip files must first be decompressed. All data layers are provided in geographic (latitude/longitude) projection, referenced to datum WGS84. In ESRI® software this projection is defined by the geographic coordinate system GCS_WGS_1984 and datum D_WGS_1984 (EPSG: 4326).<b>4. License and citations</b><b>4.1 License agreement</b>This documentation and datasets are licensed under a Creative Commons Attribution-ShareAlike 4.0 International License (CC-BY-4.0 License). For all regulations regarding license grants, copyright, redistribution restrictions, required attributions, disclaimer of warranty, indemnification, liability, waiver of damages, and a precise definition of licensed materials, please refer to the License Agreement (https://creativecommons.org/licenses/by/4.0/legalcode). For a human-readable summary of the license, please see https://creativecommons.org/licenses/by/4.0/.<b>4.2 Citations and acknowledgements.</b>Citations and acknowledgements of this dataset should be made as follows:Ehalt Macedo, H., Lehner, B., Nicell, J., Grill, G. (2024). HydroFATE (v1): A high-resolution contaminant fate model for the global river system.We kindly ask users to cite this study in any published material produced using it. If possible, online links to this repository (https://doi.org/10.6084/m9.figshare.23646282) should also be provided.<b>5. References</b>Ehalt Macedo, H., Lehner, B., Nicell, J., Grill, G. HydroFATE (v1): A high-resolution contaminant fate model for the global river system (in review), 2024.Ehalt Macedo, H., Lehner, B., Nicell, J., Grill, G., Li, J., Limtong, A., and Shakya, R.: Distribution and characteristics of wastewater treatment plants within the global river network, Earth Syst. Sci. Data, 14, 559-577, doi: 10.5194/essd-14-559-2022, 2022.Grill, G., Li, J., Khan, U., Zhong, Y., Lehner, B., Nicell, J., and Ariwi, J.: Estimating the eco-toxicological risk of estrogens in China's rivers using a high-resolution contaminant fate model, Water Research, 145, 707-720, doi: 10.1016/j.watres.2018.08.053, 2018.Klein, E. Y., Boeckel, T. P. V., Martinez, E. M., Pant, S., Gandra, S., Levin, S. A., Goossens, H., and Laxminarayan, R.: Global increase and geographic convergence in antibiotic consumption between 2000 and 2015, Proceedings of the National Academy of Sciences, 115, E3463-E3470, doi: 10.1073/pnas.1717295115, 2018.Lehner, B. and Grill, G.: Global river hydrography and network routing: baseline data and new approaches to study the world's large river systems, Hydrol Process, 27, 2171-2186, doi: 10.1002/hyp.9740, 2013.Linke, S., Lehner, B., Ouellet Dallaire, C., Ariwi, J., Grill, G., Anand, M., Beames, P., Burchard-Levine, V., Maxwell, S., Moidu, H., Tan, F., and Thieme, M.: Global hydro-environmental sub-basin and river reach characteristics at high spatial resolution, Scientific Data, 6, 283, doi: 10.1038/s41597-019-0300-6, 2019.Messager, M. L., Lehner, B., Grill, G., Nedeva, I., and Schmitt, O.: Estimating the volume and age of water stored in global lakes using a geo-statistical approach, Nature Communications, 7, 13603, doi: 10.1038/ncomms13603, 2016.<br><br><br>
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.974 | 0.978 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".