Another Step Toward “Big” Catchment Science
Notice bibliographique
Résumé
The MacroSheds project aims to catalyze large-scale and ongoing synthetic research by watershed ecosystem scientists, with the goal of developing theories that generalize across spatial scales, within and between watersheds/catchments (McDonnell et al. 2007). The heart of the project is the MacroSheds dataset (Vlah et al. 2023b), currently harmonizing streamflow, precipitation, and chemistry data from the Long-Term Ecological Research program (LTER), the Critical Zone Network of observatories (CZO/CZ Net), the U.S. Forest Service (USFS), and other programs from the national to the municipal. It provides comprehensive watershed summary statistics wherever gridded products are available. These include descriptions of terrain, vegetation, land use, soil and bedrock type, and climate: everything one needs in order to compare, categorize, and build on the results of hundreds of watershed studies, some of which have been monitoring since the 1950s (e.g., Niwot Ridge, Andrews, Hubbard Brook, Fernow LTER sites). Watershed ecosystem science began in the late 1960s, when Herb Bormann and Gene Likens began estimating precipitation inputs and streamwater exports for small gauged watersheds in the Hubbard Brook Experimental Forest (Bormann et al. 1968). These input and output fluxes and their differences were used to detect trends in air pollution, climate, rates of chemical weathering, nutrient limitation, and nutrient saturation, as well as to detect the magnitude, duration, and severity of disturbance on ecosystem element retention and loss. The simplicity of approach and magnitude of scientific impact led to watershed ecosystem studies being conducted in thousands of watersheds around the world. The era of Big Data is upon us, but not uniformly. Industry, government, and large non-governmental organization data collection efforts benefit from the standardization and aggregation that come with centralized data norms, but many academic research domains have had to pool their data resources more organically, borrowing from individually led monitoring efforts and increasingly prevalent open data initiatives. Catchment sciences straddle the domains of hydrology, geochemistry, biology, and climatology; consequently, a catchment scientist might make use of global atmospheric models and nationally managed streamflow data, while being constrained to a regional or even local scale in terms of stream chemistry data. But that picture is changing. Several recent initiatives address the need for harmonized streamflow and water quality data. GEMStat was developed in the early 2000s as a global inland water quality database and information system and remains in operation today with data from over 17,000 stations (Barker et al. 2007). In the 2010s, GEMStat was joined by the GLObal RIver Chemistry Database (GLORICH; Hartmann et al. 2014), and similar initiatives at the national or continental level. River chemistry data from three such initiatives, CESI (Canada), Waterbase (Europe), and WQP (U.S.A.; Read et al. 2017), along with both GEMStat and GLORICH, were harmonized into the Global River Water Quality Archive in 2021 (Virro et al. 2021). These efforts demonstrate a shift in capacity for broad-scale quality assessment, and some of them include watershed descriptors and flow observations, but none provides the full set of data components required for watershed chemical budgeting: precipitation and deposition, discharge and concentration, and descriptive watershed attributes. CAMELS-Chem, introduced in 2022, provides just such a collection for a subset of rivers gauged by the U.S. Geological Survey (USGS), by supplementing the existing CAMELS dataset (Addor et al. 2017; Sterle et al. 2022). MacroSheds addresses this same need, but with a focus on long-term watershed ecosystem studies, generally conducted in watersheds smaller than 10 km2, and from which accurate solute fluxes can be determined (Fig. 1). In addition to supplying high quality data, the MacroSheds project aims to lower barriers to data use. The dashboard at macrosheds.org provides flexible tools for visual exploration of the dataset, and the macrosheds R package (github.com/MacroSHEDS/macrosheds) simplifies access, analysis, and proper attribution of primary sources. The full dataset (Vlah et al. 2023b), including rich metadata, is archived through the Environmental Data Initiative, with a planned update release every January. All project code is on GitHub (github.com/MacroSHEDS), and we welcome community contributions. In the next year, we will harmonize data from additional primary sources, including the National Ecological Observatory Network, and publish monthly and annual load estimates for all chemical constituents. We will also continue development of a community upload portal, complete with both automated and manual-visual quality control. For a more thorough description of these components, please see our open-access data paper (Vlah et al. 2023a). EB and MR originated the project and defined its scope and goals. MV, MR, and SR designed the data processing system architecture. MV, SR, WS, and NG developed the data processing system, with routine feedback from MR, EB, and all other authors. Visualizations associated with the data paper and the MacroSheds portal were also designed by the full team, and generated by MV and SR. SR, MV, and WS implemented the macrosheds R package. MV, EB, MR, and SR wrote the data paper, with edits from the team.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,002 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».