BirdVox-14SD: a dataset of flight calls with species annotation
Bibliographic record
Abstract
BirdVox 14 Species Dataset (BirdVox-14SD) ============= Version 1.0, May 2020. Created By ---------- Vincent Lostanlen (1, 2, 3), Andrew Farnsworth (1), Jason Cramer (2, 3), and Juan Pablo Bello (2, 3). (1): Cornell Lab of Ornithology (CLO) (2): Center for Urban Science and Progress, New York University (3): Music and Audio Research Lab, New York University https://wp.nyu.edu/birdvox Description ----------- The BirdVox 14 Species Dataset (BirdVox-14SD) contains 14,336 audio clips of avian flight calls, each ranging from about 150 ms to 500 ms in duration. These recordings come from ROBIN autonomous recording units, placed near Ithaca, NY, USA during the 2015 migration season (August - November). Nine different sensors were used, originally numbered 1, 2, 3, 4, 5, 6, 7, 8, and 10. These sensors acquired audio recordings in intervals of two hours across the season. A subsample of 150 these two-hour recordings were chosen for annotation, using the Entrofy library [3] in order to maximize diversity across sensor locations, time of day, week in the season, and background noise characteristics (as represented by vector quantizations of median MFCCs). Andrew Farnsworth used the Raven software to pinpoint every avian flight call and labeled the corresponding order, family, and species. The dataset can be used, among other things, for the research, development and testing of bioacoustic classification models, including the reproduction of the results reported in [1]. For details on the hardware of ROBIN recording units, we refer the reader to [2]. [1] J. Cramer, V. Lostanlen, A. Farnsworth, J. Salamon, J.P. Bello. Chirping up the Right Tree: Incorporating Biological Taxonomies into Deep Bioacoustic Classifiers, Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2020. [2] J. Salamon, J. P. Bello, A. Farnsworth, M. Robbins, S. Keen, H. Klinck, and S. Kelling. Towards the Automatic Classification of Avian Flight Calls for Bioacoustic Monitoring. PLoS One, 2016. [3] D. Huppenkothen, B. McFee, L. Norén. Entrofy Your Cohort: A Data Science Approach to Candidate Selection. PLoS One, 2020. Taxonomic Annotations ----------------------- Classification annotations for each flight call are given at three taxonomic levels: order, family, and species. These annotations are condensed into a three-number-code which largely follow " . . ". The specific numeric codes are: * Order * 1.\*.\* - Passerine * Family * 1.1.\* - American Sparrow * 1.2.\* - Cardinals * 1.3.\* - Thrushes * 1.4.\* - New World warblers * Species * 1.1.1 - American tree sparrows (ATSP) * 1.1.2 - Chipping sparrow (CHSP) * 1.1.3 - Savannah sparrow (SAVS) * 1.1.4 - White-throated sparrow (WTSP) * 1.2.1 - Rose-breasted grosbeak (RBGR) * 1.3.1 - Gray-cheeked thrush (GCTH) * 1.3.2 - Swainson's thrush (SWTH) * 1.4.1 - American redstart (AMRE) * 1.4.2 - Bay-breasted warbler (BBWA) * 1.4.3 - Black-throated blue warbler (BTBW) * 1.4.4 - Canada warbler (CAWA) * 1.4.5 - Common yellowthroat (COYE) * 1.4.6 - Mourning warbler (MOWA) * 1.4.7 - Ovenbird (OVEN) Additionally, at any level of the taxonomy, the numeric code "0" is reserved for "other" and the code "X" refers to unknown. For example, 1.1.0 corresponds to an American Sparrow with a species outside of our scope of interest, and 1.1.X corresponds to an American Sparrow of unknown species. At the top level (family), the "other" codes (0.\*.\*) deviate from the family-order-species in order to capture a variety of other out-of-scope sounds, including anthropophony, non-avian biophony, and biophony of avians outside of the scope of interest. The file `taxonomy.yaml` details this taxonomy structure. Data Files ------------ BirdVox-14SD contains the recordings as HDF5 files, sampled at 22,050 Hz, with a single channel (mono). Each HDF5 file contains flight call vocalizations of a particular species. The name of each HDF5 file follows the format: `BirdVox-14SD_ _original.h5`. The name of the HDF5 dataset in each file is "waveforms", with the corresponding key for each audio recording following the format: `unit - _ _ `. Please acknowledge BirdVox-14SD in academic research -------------------------------------------------------------------------- When BirdVox-14SD is used for academic research, we would highly appreciate it if scientific publications of works partly based on this dataset cite the following publication: J. Cramer, V. Lostanlen, A. Farnsworth, J. Salamon, J.P. Bello. Chirping up the Right Tree: Incorporating Biological Taxonomies into Deep Bioacoustic Classifiers, Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2020. The creation of this dataset was supported by NSF grants 1633259 (BIRDVOX). Conditions of Use ---------------------- Dataset created by Vincent Lostanlen, Andrew Farnsworth, Jason Cramer, Justin Salamon, and Juan Pablo Bello. The BirdVox-14SD dataset is offered free of charge under the terms of the Creative Commons Attribution 4.0 International License. The dataset and its contents are made available on an "as is" basis and without warranties of any kind, including without limitation satisfactory quality and conformity, merchantability, fitness for a particular purpose, accuracy or completeness, or absence of errors. Subject to any liability that may not be excluded or limited by law, CLO is not liable for, and expressly excludes all liability for, loss or damage however and whenever caused to anyone by any use of the BirdVox-14SD dataset or any part of it. Feedback ----------- Please help us improve BirdVox-14SD by sending your feedback to: vincent.lostanlen@gmail.com and jtcramer@nyu.edu In case of a problem, please include as many details as possible. Acknowledgements ------------------------ Jessie Barry, Ian Davies, Tom Fredericks, Jeff Gerbracht, Sara Keen, Holger Klinck, Anne Klingensmith, Ray Mack, Peter Marchetto, Ed Moore, Matt Robbins, Ken Rosenberg, and Chris Tessaglia-Hymes. We acknowledge that the land on which the data was collected is the unceded territory of the Cayuga nation, which is part of the Haudenosaunee (Iroquois) confederacy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.021 | 0.040 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".