<scp>ATLANTIC</scp>‐<scp>PRIMATES</scp>: a dataset of communities and occurrences of primates in the Atlantic Forests of South America
Bibliographic record
Abstract
Abstract Primates play an important role in ecosystem functioning and offer critical insights into human evolution, biology, behavior, and emerging infectious diseases. There are 26 primate species in the Atlantic Forests of South America, 19 of them endemic. We compiled a dataset of 5,472 georeferenced locations of 26 native and 1 introduced primate species, as hybrids in the genera Callithrix and Alouatta. The dataset includes 700 primate communities, 8,121 single species occurrences and 714 estimates of primate population sizes, covering most natural forest types of the tropical and subtropical Atlantic Forest of Brazil, Paraguay and Argentina and some other biomes. On average, primate communities of the Atlantic Forest harbor 2 ± 1 species (range = 1–6). However, about 40% of primate communities contain only one species. Alouatta guariba (N = 2,188 records) and Sapajus nigritus (N = 1,127) were the species with the most records. Callicebus barbarabrownae (N = 35), Leontopithecus caissara (N = 38), and Sapajus libidinosus (N = 41) were the species with the least records. Recorded primate densities varied from 0.004 individuals/km2 (Alouatta guariba at Fragmento do Bugre, Paraná, Brazil) to 400 individuals/km2 (Alouatta caraya in Santiago, Rio Grande do Sul, Brazil). Our dataset reflects disparity between the numerous primate census conducted in the Atlantic Forest, in contrast to the scarcity of estimates of population sizes and densities. With these data, researchers can develop different macroecological and regional level studies, focusing on communities, populations, species co‐occurrence and distribution patterns. Moreover, the data can also be used to assess the consequences of fragmentation, defaunation, and disease outbreaks on different ecological processes, such as trophic cascades, species invasion or extinction, and community dynamics. There are no copyright restrictions. Please cite this Data Paper when the data are used in publications. We also request that researchers and teachers inform us of how they are using the data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".