Stenert_et_al_GCB_gcb.15367_Data_repository
Bibliographic record
Abstract
These datasets contain macroinvertebrate sampling data for depressional wetlands from several study regions across North and South America. They represent the non-federally funded data supporting the results of the primary research paper entitled "Climate‐ versus geographic‐dependent patterns in the spatial distribution of macroinvertebrate assemblages in New World depressional wetlands", authored by Stenert et al. (Global Change Biology, 2020; https://doi.org/10.1111/gcb.15367). These datasets were employed in a two-part analytical approach that assessed patterns in the macroinvertebrate assemblages of depressional wetlands across the temperate and subtropical climatic zones of North and South America. We aimed to better understand how wetland macroinvertebrates assemblages were structured according to geography and climate. To do so, we contrasted aquatic‐macroinvertebrate assemblage structure (family‐level) between subtropical and temperate depressional wetlands of North and South America using presence‐absence data from 264 of these habitats across the continents and more‐detailed relative‐abundance data from 56 depressional wetlands from four case study locations (North Dakota and Georgia in North America; southern Brazil and Argentinian Patagonia in South America). The dataset "Intercontinental analysis" includes 127 wetlands from seven study regions in North America covering subtropical (N = 32; USA states of Georgia, New Mexico, South Carolina and Texas) and temperate climates (N = 78; USA states of Iowa, Michigan, Minnesota and Wisconsin; Canada province of Ontario). The dataset from South America included 137 wetlands from two study regions covering subtropical (N = 72; Brazil state of Rio Grande do Sul) and temperate climates (N = 65; Argentinean Patagonia; provinces of Chubut, Santa Cruz and Tierra del Fuego). The dataset "Case-study analysis" includes data from three specific locations in temperate and subtropical climate zones of North and South America. The subtropical wetlands were located in the Southeastern USA (state of Georgia) and Southern Brazil (state of Rio Grande do Sul). The temperate wetlands were located in Argentinean Patagonia. We used ordination methods (PCA and NMDS) and tests of multivariate dispersion (PERMDISP) to assess the distribution and the homogeneity in variation in the composition of macroinvertebrate assemblages across climates and continents, respectively. Taxonomic identification was conducted to the family level, except for planarians, water mites and some Anostraca and Oligochaeta, which were left at the lowest taxonomic level practical. Bryozoa (Plumatellidae), Cnidaria (Hydridae), Platyhelminthes (Turbellaria), Annelida (Clitellata: Oligochaeta and Hirudinea), Mollusca (Bivalvia and Gastropoda) and Arthropoda (Crustacea: Branchiopoda and Malacostraca; Arachnida; Insecta) were the phyla (and their corresponding subphyla and/or classes) considered in this study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.008 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.012 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.240 | 0.138 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".