Manual versus automatic identification of black-capped chickadee (Poecile atricapillus) vocalizations
Bibliographic record
Abstract
One time-consuming aspect of bioacoustic research is identifying vocalizations from long audio recordings. SongScope (version 4.1.5. Wildlife Acoustics, Inc.) is a computer program capable of developing acoustic recognizers that can identify wildlife vocalizations. The goal of the current study was to compare the effectiveness of manual identification of black-capped chickadee vocalizations to identification by SongScope recognizers. A recognizer was developed for each main chickadee vocalization by providing previously annotated audio of chickadees. Six chickadees (three male, three female) were recorded in one-hour intervals with and without anthropogenic (i.e., man-made) noise to provide a variety of samples to test the recognizer. These recordings were analyzed via the recognizer and two human coders, with an additional third coder reviewing a random subset of recordings for reliability. Strong agreement was found between the human coders, κ = 0.76, p < 0.00. Agreement between human coders and the recognizer was moderate for fee songs, κ = 0.46, p < 0.00, and strong for fee-bee songs, κ = 0.77, p < 0.00, as well as for chick-a-dee calls, κ = 0.82, p < 0.00. Results showed that male chickadees produced more tseet calls in silence and females produced more gargle calls during noise. No differences were found in vocalizations based on time of day. Our observations also suggest that the chick-a-dee recognizer was capable of identifying gargle and tseet calls along with the intended chick-a-dee calls. Overall, SongScope was effective at identifying fee-bee songs and chick-a-dee calls, but not as effective for identifying fee songs. These recognizers can allow for faster acoustic analyses (by approximately four times) and be continuously improved for greater accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".