<scp>LandFrag</scp>: A Dataset to Investigate the Effects of Forest Loss and Fragmentation on Biodiversity
Bibliographic record
Abstract
ABSTRACT Motivation The accelerated and widespread conversion of once continuous ecosystems into fragmented landscapes has driven ecological research to understand the response of biodiversity to local (fragment size) and landscape (forest cover and fragmentation) changes. This information has important theoretical and applied implications, but is still far from complete. We compiled the most comprehensive and updated database to investigate how these local and landscape changes determine species composition, abundance and trait diversity of multiple taxonomic groups in forest fragments across the globe. Main Types of Variables Contained We gathered data for 1472 forest fragments, providing information on the abundance and composition of 9154 species belonging to vertebrates, invertebrates, and plants. For 2703 of these species, we obtained more than 20 functional traits. We provided the spatial location and size of each fragment and metrics of landscape composition and configuration. Spatial Location and Grain The dataset includes 1472 forest fragments sampled in 121 studies from all continents except Antarctica. Most datasets (77%) are from tropical regions, 17% are from temperate regions, and 6% are from subtropical regions. Species abundance and composition were collected at the plot or fragment scale, whereas the landscape metrics were extracted with buffer size ranging from a radius of 200–2000 m. Time Period and Grain Data on the abundance of species and community composition were collected between 1994 and 2022, and the landscape metrics were extracted from the same year that a given study collected the abundance and composition data. Major Taxa and Level of Measurement The studied organisms included invertebrates (Arachnida, Insecta and Gastropoda; 41% of the datasets), vertebrates (Amphibia, Squamata, Aves and Mammalia; 44%), and vascular plants (19%), and the lowest level of identification was species or morphospecies. Software Format The dataset and code can be downloaded on Zenodo or GitHub.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.011 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.015 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".