Estimation of avian species richness: biases in morning surveys and efficient sampling from acoustic recordings
Bibliographic record
Abstract
Abstract Species richness estimation is an important component of ecological studies and conservation planning. Limited resources necessitate that sampling protocols be as efficient and accurate as possible. For birds, automated acoustic sampling offers potential advantages of abundant data at reduced cost for field observers, and enhanced diel coverage, but neither of which may accrue if surveys are biased and/or too costly to analyze in the lab. Here, we assessed bias in estimates of species and higher order taxonomic richness obtained from standard morning point counts, and from morning‐only acoustic recordings, relative to estimates from 72, 10‐min acoustic recordings conducted hourly over 3 d. Furthermore, we compared 10‐min subsamples of 24‐h recordings across five statistical estimators to establish which combination of number of samples, from which times of day, and with which statistical estimator, best approximated total observed species richness. Total observed species richness was the total number of species detected per site over 720 min of 10‐min recordings. Standard morning point counts and morning‐only acoustic recordings consistently underestimated both total species and higher order taxonomic richness. Species not detected were those that irregularly or nocturnally vocalize. Without statistical estimators, the greatest number of species per unit sample effort was detected from 10‐min, on‐the‐hour samples between 07:00 and 12:00, and at 21:00, over 3 d. With the jackknife estimator, three 10‐min samples (one at each of 08:00, 09:00, and 12:00, over 3 d) most efficiently estimated within 5% of total observed species richness. Researchers can subsample in combination with statistical estimators to increase analytical efficiency for species richness using acoustic recordings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.017 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".