MétaCan
Menu
Back to cohort

A Comprehensive Surface Water Quality Monitoring Dataset (1940-2023): 2.82Million Record Resource for Empirical and ML-Based Research

2025· dataset· en· W6921140991 on OpenAlexaboutno aff

Bibliographic record

VenueFigshare · 2025
Typedataset
Languageen
Field
Topic
Canadian institutionsnot available
Fundersnot available
KeywordsOutlierWater qualityIndex (typography)Missing dataQuality (philosophy)Resource (disambiguation)Sample (material)Data qualityData collection

Abstract

fetched live from OpenAlex

<b>Data Description</b><b>Water Quality Parameters</b>: Ammonia, BOD, DO, Orthophosphate, pH, Temperature, Nitrogen, Nitrate.<b>Countries/Regions</b>: United States, Canada, Ireland, England, China.<b>Years Covered</b>: 1940-2023.<b>Data Records</b>: 2.82 million.<b>Definition of Columns</b><b>Country</b>: Name of the water-body region.<b>Area</b>: Name of the area in the region.<b>Waterbody Type</b>: Type of the water-body source.<b>Date</b>: Date of the sample collection (dd-mm-yyyy).<b>Ammonia (mg/l)</b>: Ammonia concentration.<b>Biochemical Oxygen Demand (BOD) (mg/l)</b>: Oxygen demand measurement.<b>Dissolved Oxygen (DO) (mg/l)</b>: Concentration of dissolved oxygen.<b>Orthophosphate (mg/l)</b>: Orthophosphate concentration.<b>pH (pH units)</b>: pH level of water.<b>Temperature (°C)</b>: Temperature in Celsius.<b>Nitrogen (mg/l)</b>: Total nitrogen concentration.<b>Nitrate (mg/l)</b>: Nitrate concentration.<b>CCME_Values</b>: Calculated water quality index values using the CCME WQI model.<b>CCME_WQI</b>: Water Quality Index classification based on CCME_Values.<b>Data Directory Description:</b><b>Category 1: Dataset</b><b>Combined Data: </b>This folder contains two CSV files: <i>Combined_dataset.csv</i> and <i>Summary.xlsx</i>. The <i>Combined_dataset.csv</i> file includes all eight water quality parameter readings across five countries, with additional data for initial preprocessing steps like missing value handling, outlier detection, and other operations. It also contains the CCME Water Quality Index calculation for empirical analysis and ML-based research. The <i>Summary.xlsx</i> provides a brief description of the datasets, including data distributions (e.g., maximum, minimum, mean, standard deviation).<br><i>Combined_dataset.csv</i><i>Summary.xlsx</i><b>Country-wise Data: </b>This folder contains separate country-based datasets in CSV files. Each file includes the eight water quality parameters for regional analysis. The <i>Summary_country.xlsx</i> file presents country-wise dataset descriptions with data distributions (e.g., maximum, minimum, mean, standard deviation).<br><i>England_dataset.csv</i><i>Canada_dataset.csv</i><i>USA_dataset.csv</i><i>Ireland_dataset.csv</i><i>China_dataset.csv</i><i>Summary_country.xlsx</i><b>Category 2: Code</b><br>Data processing and harmonization code (e.g., Language Conversion, Date Conversion, Parameter Naming and Unit Conversion, Missing Value Handling, WQI Measurement and Classification).<br><i>Data_Processing_Harmonnization.ipynb</i>The code used for Technical Validation (e.g., assessing the Data Distribution, Outlier Detection, Water Quality Trend Analysis, and Vrifying the Application of the Dataset for the ML Models).<i>Technical_Validation.ipynb</i><b>Category 3: Data Collection Sources</b><br>This category includes links to the selected dataset sources, which were used to create the dataset and are provided for further reconstruction or data formation. It contains links to various data collection sources.<br><i>DataCollectionSources.xlsx</i><b>Original Paper Title: </b>A Comprehensive Dataset of Surface Water Quality Spanning 1940-2023 for Empirical and ML Adopted Research<b>Abstract</b><br>Assessment and monitoring of surface water quality are essential for food security, public health, and ecosystem protection. Although water quality monitoring is a known phenomenon, little effort has been made to offer a comprehensive and harmonized dataset for surface water at the global scale. This study presents a comprehensive surface water quality dataset that preserves spatio-temporal variability, integrity, consistency, and depth of the data to facilitate empirical and data-driven evaluation, prediction, and forecasting. The dataset is assembled from a range of sources, including regional and global water quality databases, water management organizations, and individual research projects from five prominent countries in the world, e.g., the USA, Canada, Ireland, England, and China. The resulting dataset consists of 2.82 million measurements of eight water quality parameters that span 1940 - 2023. This dataset can support meta-analysis of water quality models and can facilitate Machine Learning (ML) based data and model-driven investigation of the spatial and temporal drivers and patterns of surface water quality at a cross-regional to global scale.<br><b>Note:</b> Cite this repository and the original paper when using this dataset.<br><br>

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.006
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.009
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.006
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0010.001
Science and technology studies0.0010.000
Scholarly communication0.0010.000
Open science0.0020.003
Research integrity0.0010.003
Insufficient payload (model declined to judge)0.0140.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.500
GPT teacher head0.511
Teacher spread0.011 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueFigshareFrench-language works237,207