Bird–window collisions: A comprehensive dataset for the Neotropical region
Bibliographic record
Abstract
Our primary objective was to compile a comprehensive dataset on bird-window collisions throughout the Neotropical region, including both published and unpublished sources. On May 12, 2020, we extensively disseminated invitations to provide data via email and social media platforms. By providing a template worksheet, we required standardized information from collaborators to complete and register their data. To better understand how these data were acquired (e.g., incidental observations and systematic procedures), we sent out a survey to all collaborators. We established rigorous validation criteria for data inclusion and conducted thorough curation procedures to ensure accuracy. After the filtering process, we compiled a total of 4103 bird-window collision reports. These came from 11 Neotropical countries, dating from 1946 to 2020, and revealing distinct regional patterns and potential seasonal patterns. The five most frequent orders were Passeriformes (2451), Columbiformes (520), Apodiformes (377), Psittaciformes (202), and Piciformes (186). Data on bird-window collisions were collected through a local specific systematic protocol (1419), by chance (1252), by government agencies (742), and by other approaches (632), while a few reports were collected by unknown procedures (58). The volume of records across months in our dataset suggests that there may be temporal patterns, with peaks: the first one in March-April and the second one in October-November, which seem to align with the major migration and reproduction seasons. This dataset represents the first comprehensive effort in the Neotropical region focused on bird-window collision data, providing valuable insights for further scientific advancements and conservation policies. The data are free from copyright or proprietary restrictions. Please cite this data paper when using the data in publications or scientific presentations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".