Code and data: Ecological stability propagates across spatial scales and trophic levels in freshwater ecosystems
Bibliographic record
Abstract
The files listed below include data and code used to produce the results, including those made available as supplementary material, described in Siqueira et al. (https://ecoevorxiv.org/mpf5x), "Ecological stability propagates across spatial scales and trophic levels in freshwater ecosystems". The full set of results can be reproduced by running four R codes following this sequence: - 01_Siqueira_etal_dataprep_stability_metrics.R - 02_Siqueira_etal_SEM_analyses.R - 03_Siqueira_etal_stab_figs.R - 04_Siqueira_etal_stab_supp_m.R To run the first R code (“01_Siqueira_etal_dataprep_stability_metrics.R”), 3 .csv files are needed. "data_stab_analysis_except_lepas.csv" has 151404 rows and 17 columns. Rows represent the abundance (individual counts, biomass, coverage area) of a given taxon, at a given site, in a given year. Some abundance values were generated by interpolation (see details in the Methods section of the paper and supplementary material). The 17 columns include "Data_set_ID": a numerical vector describing 34 of the 35 independent metacommunity data sets. For the 35th data set, see below. "Metacom": a character vector describing 34 of the 35 independent metacommunity data sets. This column was kept for the sake of preserving the original identity of each data set. "Sample_id": a vector representing a unique combination of data set ID, site ID, and time step. "Site": a vector describing the ID of the site where the sample was taken. "Year": a vector describing the year when the sample was taken. "Frequency": a vector describing the frequency interval at which the samples were taken. This variable has only one state (“Inter_annual”) and was kept because some of the data originally gathered (before filtering) included also intra-annual sampling (these were discarded). "Time_step": a vector describing the time step a sample was taken, considering the whole number of years included in that data set. "Lat": a vector describing the latitude of the sampling site in decimal degrees. "Long": a vector describing the latitude of the sampling site in decimal degrees. "Datum": a vector describing the datum used by the person who collected the sample. "Coord_system": a vector describing the coordinate system used in this data set after data cleaning. "Bio_gr": a vector describing the major biological groups in the data set (e.g., macroinvertebrates, fish). "Ecosys": a vector describing the two ecosystem types in the data set (streams and lakes). "Species": a vector describing the species or genus names. "Abundance": a vector describing species or genus abundance, density, or coverage. "Trophic_gr": a vector describing 3 broad trophic groups (producers, consumers that are invertebrates, and consumers that are vertebrates). "New_tr_g": a vector describing 5 trophic levels used in the analysis (producers, primary consumers secondary consumers, tertiary consumers). The following files were made available separately due to the data sharing policies of The Ohio Division of Wildlife (ODOW). Thus, we made them available not in the format of raw numbers, but as pre-processed variables that we used in our models (e.g., stability and diversity metrics, trophic groups, etc.). "lepas_site_stab_metrics.csv" has 8 rows and 14 columns. Rows represent sites sampled for a given trophic level. Columns include "Lat", "Long", "Metacom", "Sample_id", "Site", "Bio_gr", "Ecosys", and "New_tr_g" as described above and: "Site_troph": a vector representing a combination of site ID and trophic level. This was used to identify sites that include more than one trophic level. "Simp": a vector describing the Simpson diversity index of a sampled site. "synchrony_comm_site": a vector describing species population synchrony within a sampled site. "mean_cv_species_site": a vector describing the average coefficient of variation of species sampled in a site. "cv_comm_site": a vector describing local temporal variability of aggregated community abundance. "S": a vector describing species richness in a sampled site. "lepas_meta_stab_metrics.csv" has 1 row and 21 columns. The row represents the LEPAS metacommunity. Columns include "Metacom", "Freq", "Time_step", "Bio_gr", "Trophic_gr", "New_tr_g", and "Ecosys" as described above and: "Meta_troph": a vector representing a combination of metacommunity ID and trophic level. This was used to differentiate metacommunities that include more than one trophic level. "Simp": a vector describing the Simpson diversity index averaged across sites within the metacommunity. "S": a vector describing the species richness averaged across sites within the metacommunity. "nSites": a vector describing the number of sampled sites in the metacommunity. "CV_S_L": a vector describing population temporal variability (sensu Wang et al. 2019) within metacommunity. "CV_C_L": a vector describing community temporal variability (sensu Wang et al. 2019) within metacommunity. "CV_S_R": a vector describing metapopulation temporal variability (sensu Wang et al. 2019) within metacommunity. "CV_C_R": a vector describing metacommunity temporal variability (sensu Wang et al. 2019). "phi_S_L2R": a vector describing species-level spatial synchrony (sensu Wang et al. 2019) within metacommunity. "phi_C_L2R": a vector describing community-level spatial synchrony (sensu Wang et al. 2019) within metacommunity. "phi_S2C_L": a vector describing local-scale species synchrony (sensu Wang et al. 2019) within metacommunity. "phi_S2C_R": a vector describing regional-scale species synchrony (sensu Wang et al. 2019) within metacommunity. "Simp_gamma": a vector describing Simpson gamma diversity in the metacommunity. "S_gamma": a vector describing regional richness in the metacommunity. To run the second R code (“02_Siqueira_etal_SEM_analyses.R”), two groups of files are needed – one group of files prepared with the first code and the following two .csv files: "site_env_preds.csv" has 735 rows and 3 columns. Rows represent climatic data extracted from each of the sites used in analyses. The 3 columns include: "Site_troph": a vector representing a combination of site ID and trophic level. This was used to identify sites that include more than one trophic level. "bio4": a numerical vector describing seasonality in temperature in each site. "bio15": a numerical vector describing seasonality in precipitation in each site. "meta_env_spa_preds.cs" has 59 rows and 4 columns. Rows represent each of the sites used in analyses. The 4 columns include: "Meta_troph": a vector representing a combination of metacommunity ID and trophic level. This was used to differentiate metacommunities that include more than one trophic level. "Tmax_sync ": a numerical vector describing synchrony in maximum temperature within the metacommunity. "Tmin_sync ": a numerical vector describing synchrony in minimum temperature within the metacommunity. "Precip_sync ": a numerical vector describing synchrony in precipitation within the metacommunity. To run the third and fourth R codes (“03_Siqueira_etal_stab_figs.R”; “04_Siqueira_etal_stab_supp_m.R”), one needs the same files described above.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.045 |
| Meta-epidemiology (narrow) | 0.004 | 0.003 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.390 | 0.227 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".