Compiled multi-source suspended sediment concentration in-situ data
Bibliographic record
Abstract
The curated dataset Compiled multi-source suspended sediment concentration in-situ data is a global compilation of several datasets (listed below). When using this dataset, please cite the curators' preprint / published paper, if available, as well as the original data sources. Citing all the original data sources is imperative. Curator reference: (preprint) Luisa Vieira Lucchese, Rangel Daroya, Travis T Simmons, et al. Operational near real time global riverine sediment flux estimates from space. ESS Open Archive . November 17, 2025. DOI: 10.22541/essoar.176339882.25879825/v1 Original data sources references: Pimenta et al., 2020: Pimenta, A., McKinney, R., Hanson, A., Johnson, R., Cobb, D., Janiec, J., & Oczkowski, A. (2020). Water quality data for Narragansett Bay, RI (USA) from 2014 to 2018 [Data set]. Zenodo. https://doi.org/10.5281/zenodo.3747391 MCN-WQM: Misipawistik Cree Nation. 2023. "MCN-WQM" (dataset). 1.0.0. DataStream. https://doi.org/10.25976/uhjn-1746. Dethier et al., 2020: https://doi.org/10.1029/2019JF005033 and Dethier, E. (2019). Suspended sediment concentration database, HydroShare, http://www.hydroshare.org/resource/2ee7d421618a4873b9906540d047ced4, they reference the original data sources:ANA (2018), Agência Nacional de Águas, edited, July, 2017.HYDAT (2018), The Water Survey of Canada, edited, August, 2018.USGS (2018), U.S. Geological Survey, edited, August, 2018.WRA (2018), Taiwan Water Resource Agency, edited, August, 2018. GEMStat: GEMStat - The global freshwater quality database. (n.d.). Retrieved October 11, 2023, from https://gemstat.org HyBam: SO-HyBam – Service d’observation des ressources en eaux du bassin de l’Amazone. (n.d.). Retrieved October 11, 2023, from https://hybam.obs-mip.fr/ Zhai et al., 2024: Zhai, M., Zhou, X., Tao, Z., Xie, Y., Yang, J., Shao, W., Zhang, H., & Lv, T. (2024). Satellite-ground synchronous in-situ dataset of water optical parameters and surface temperature for typical lakes in China [Data set]. In Scientific Data. Zenodo. https://doi.org/10.5281/zenodo.13777017 and Zhai, M., Zhou, X., Tao, Z. et al. Satellite-ground synchronous in-situ dataset of water optical parameters and surface temperature for typical lakes in China. Sci Data 11, 883 (2024). https://doi.org/10.1038/s41597-024-03704-3 Van der Hiele, 2021: Van der Hiele, T. (2021). Dataset of sampled and/or logged Chlorophyll, Total Suspended Matter, Turbidity and Water temperature in Dutch Case study areas [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5121554 Waterbase: 12. European Environment Agency. (2022). Waterbase [Data set]. Retrieved from https://www.eea.europa.eu/en/datahub/datahubitem-view/208518d1-ffe3-4981-9cae-13264cd9c32c To check the source of a given sample, consult the Database column of the .csv file. The WaterType column indicates the water types for each sample, available as metadata in the original samples and/or after retrieval through the dataRetrieval API of USGS (https://doi-usgs.github.io/dataRetrieval/).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.026 | 0.039 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".