MétaCan
Menu
Back to cohort
Record W4385358369 · doi:10.1002/lob.10590

Another Step Toward “Big” Catchment Science

2023· article· en· W4385358369 on OpenAlexaboutno aff
Michael Vlah, Emily S. Bernhardt, Spencer Rhea, Weston Slaughter, Amanda Delvecchia, Audrey Thellman, Matthew Ross

Bibliographic record

VenueLimnology and Oceanography Bulletin · 2023
Typearticle
Languageen
FieldEnvironmental Science
TopicHydrology and Watershed Management Studies
Canadian institutionsnot available
FundersNational Science Foundation
KeywordsChapelState (computer science)Library scienceArt historyHistoryComputer science

Abstract

fetched live from OpenAlex

The MacroSheds project aims to catalyze large-scale and ongoing synthetic research by watershed ecosystem scientists, with the goal of developing theories that generalize across spatial scales, within and between watersheds/catchments (McDonnell et al. 2007). The heart of the project is the MacroSheds dataset (Vlah et al. 2023b), currently harmonizing streamflow, precipitation, and chemistry data from the Long-Term Ecological Research program (LTER), the Critical Zone Network of observatories (CZO/CZ Net), the U.S. Forest Service (USFS), and other programs from the national to the municipal. It provides comprehensive watershed summary statistics wherever gridded products are available. These include descriptions of terrain, vegetation, land use, soil and bedrock type, and climate: everything one needs in order to compare, categorize, and build on the results of hundreds of watershed studies, some of which have been monitoring since the 1950s (e.g., Niwot Ridge, Andrews, Hubbard Brook, Fernow LTER sites). Watershed ecosystem science began in the late 1960s, when Herb Bormann and Gene Likens began estimating precipitation inputs and streamwater exports for small gauged watersheds in the Hubbard Brook Experimental Forest (Bormann et al. 1968). These input and output fluxes and their differences were used to detect trends in air pollution, climate, rates of chemical weathering, nutrient limitation, and nutrient saturation, as well as to detect the magnitude, duration, and severity of disturbance on ecosystem element retention and loss. The simplicity of approach and magnitude of scientific impact led to watershed ecosystem studies being conducted in thousands of watersheds around the world. The era of Big Data is upon us, but not uniformly. Industry, government, and large non-governmental organization data collection efforts benefit from the standardization and aggregation that come with centralized data norms, but many academic research domains have had to pool their data resources more organically, borrowing from individually led monitoring efforts and increasingly prevalent open data initiatives. Catchment sciences straddle the domains of hydrology, geochemistry, biology, and climatology; consequently, a catchment scientist might make use of global atmospheric models and nationally managed streamflow data, while being constrained to a regional or even local scale in terms of stream chemistry data. But that picture is changing. Several recent initiatives address the need for harmonized streamflow and water quality data. GEMStat was developed in the early 2000s as a global inland water quality database and information system and remains in operation today with data from over 17,000 stations (Barker et al. 2007). In the 2010s, GEMStat was joined by the GLObal RIver Chemistry Database (GLORICH; Hartmann et al. 2014), and similar initiatives at the national or continental level. River chemistry data from three such initiatives, CESI (Canada), Waterbase (Europe), and WQP (U.S.A.; Read et al. 2017), along with both GEMStat and GLORICH, were harmonized into the Global River Water Quality Archive in 2021 (Virro et al. 2021). These efforts demonstrate a shift in capacity for broad-scale quality assessment, and some of them include watershed descriptors and flow observations, but none provides the full set of data components required for watershed chemical budgeting: precipitation and deposition, discharge and concentration, and descriptive watershed attributes. CAMELS-Chem, introduced in 2022, provides just such a collection for a subset of rivers gauged by the U.S. Geological Survey (USGS), by supplementing the existing CAMELS dataset (Addor et al. 2017; Sterle et al. 2022). MacroSheds addresses this same need, but with a focus on long-term watershed ecosystem studies, generally conducted in watersheds smaller than 10 km2, and from which accurate solute fluxes can be determined (Fig. 1). In addition to supplying high quality data, the MacroSheds project aims to lower barriers to data use. The dashboard at macrosheds.org provides flexible tools for visual exploration of the dataset, and the macrosheds R package (github.com/MacroSHEDS/macrosheds) simplifies access, analysis, and proper attribution of primary sources. The full dataset (Vlah et al. 2023b), including rich metadata, is archived through the Environmental Data Initiative, with a planned update release every January. All project code is on GitHub (github.com/MacroSHEDS), and we welcome community contributions. In the next year, we will harmonize data from additional primary sources, including the National Ecological Observatory Network, and publish monthly and annual load estimates for all chemical constituents. We will also continue development of a community upload portal, complete with both automated and manual-visual quality control. For a more thorough description of these components, please see our open-access data paper (Vlah et al. 2023a). EB and MR originated the project and defined its scope and goals. MV, MR, and SR designed the data processing system architecture. MV, SR, WS, and NG developed the data processing system, with routine feedback from MR, EB, and all other authors. Visualizations associated with the data paper and the MacroSheds portal were also designed by the full team, and generated by MV and SR. SR, MV, and WS implemented the macrosheds R package. MV, EB, MR, and SR wrote the data paper, with edits from the team.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.311
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.002
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.013
GPT teacher head0.216
Teacher spread0.203 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueLimnology and Oceanography BulletinSame topicHydrology and Watershed Management StudiesFrench-language works237,207