HydroClim Data Portal: Cyberinfrastructure for providing high-resolution GIS modeled streamflow and water temperature data to researchers
Bibliographic record
Abstract
Freshwater ecosystems play a key role in sustaining aquatic biodiversity. However, human alterations to watersheds and climate change are reducing critical habitat and the viability of populations of many aquatic species. The environmental changes have also had significant adverse impacts on water temperatures and streamflow. The changes in temperature and precipitation forecast over the next century are expected to affect the freshwater ecosystems and their biodiversity to an even greater extent than in the past. The aims of the HydroClim project are to provide openly accessible data on two key measures of stream conditions in the United States (US) and Canada for use in research, to increase public understanding of issues involving water resources, and to provide training opportunities for scientists who will be responsible for the conservation of freshwater biodiversity in the future. The project has used contemporary air temperature and precipitation data and future climate data from multiple Global Climate Model scenarios to generated high-resolution, spatially explicit, monthly streamflow and water temperature data for all watersheds across the US and Canada from 1950–2099 through multiple Soil and Water Assessment Tool (SWAT) hydrologic models. This presentation describes a cyberinfrastructure we developed for hosting the HydroClim data, consisting of a relational database and a web-based data portal that allows scientists to query and download the data. We have imported almost 1.9 billion HydroClim data records into the system. At the time of this submission, 1.3 billion records of historical data and predicted streamflow and water temperature model data are available in the HydroClim data portal for 26 watersheds in the United States. The HydroClim data are also being integrated with fish occurrence data from Fishnet 2, via the Fishnet 2 API (Application Programming Interface), which provides occurence data records for over 4.1 million species lots representing over 40 million specimens in ichthyological research collections. Our plan is to extract and merge environmental data from Hydroclim API, with fish occurrences containing geospatial information from the Fishnet 2 API, displaying the integrated data on web-based interactive hydrological maps in different time-series, and providing a tool for visualizing ecosystem diversity. The combined Hydroclim and Fishnet2 data can be used for ecological niche modeling applications, such as predicting the future distribution of threatened and endangered freshwater fish species. I will describe the cyberinfrastructure of HydroClim data portal and some of the ways the data can be used in biodiverisity research in the future.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.005 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".