Effect of sampling effort on bias and precision of trends in migration counts
Bibliographic record
Abstract
Hourly or daily counts of animals during migration are used to assess change in population status. However, due to financial and logistical constraints, it is sometimes not possible to sample the entire migration, resulting in an unknown impact on accuracy and precision of estimated trends. Using simulated migration counts of 3 raptor species, commonly detected Sharp-shinned Hawks (Accipiter striatus), rarely detected Merlins (Falco columbarius), and super-flocking Broad-winged Hawks (Buteo platypterus), we used a hierarchical modeling framework to test whether sampling effort influenced accuracy and precision of trends. We subsampled simulated datasets in various ways: weekends only; a random sample of 40%, 60%, or 80% of the migration; or the first 25%, 50%, or 75% of the migration. Bias in trend estimates and probability of error (estimating a false but significant trend) varied little with sampling effort for the common and rare species, which suggests that for these species, sampling only a small subset of the migration will not increase the probability of drawing false inference from the data. However, when counts of the super-flocking species were highly variable among days, trends were strongly positively biased when the first 25% of the migration was sampled, and probability of error increased with sampling effort. For all species, probability of detecting a significant and accurate trend also improved with sampling effort, but only exceeded 80% for common and rare species when a majority of the migration was sampled. Increasing sampling effort is recommended to improve power of trend analyses for common and rare species, but for super-flocking species, randomly sampling a subset of observation days throughout the migration can minimize error, but at the expense of a ~22–25% reduction in power using 20-year datasets. For species with this pattern, other methods to improve power, such as combining data across sites, should be considered.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".