Population-level migration modeling of North America’s birds through data integration with BirdFlow
Bibliographic record
Abstract
Abstract Background Accurate information on population-level movements of migratory animals is essential for understanding migration and for designing effective conservation strategies in a changing world. Yet such information remains scarce for most migratory species due to the effort and expense needed to collect data across their full distribution ranges. BirdFlow is a probabilistic modeling framework that infers population-level movements from weekly species distribution maps produced by the participatory science project eBird. However, BirdFlow models have only been tuned for a handful of species using high-resolution individual tracking data, which is not available for most migratory species. Methods Here, we introduce a general tuning and evaluation framework for BirdFlow that enables the first large-scale integration of distributional and individual-level data to infer animal movement across continents and hundreds of migratory species, eliminating reliance on any single individual-tracking data source. By generalizing the BirdFlow model parametrization, we enable tuning and validation using multiple complementary data sources, including GPS tracks, banding recoveries, and radio telemetry data from the Motus Wildlife Tracking System. We investigate the efficacy of this approach by (1) investigating predictive performance compared to null models; (2) validating the biological plausibility of BirdFlow models by comparing movement properties such as route straightness, number of stopovers, and migration speed between model-generated routes and real movement tracks; and (3) comparing the performance of models tuned on species-specific movement data to models tuned using hyperparameters transferred from other species. Results Our results show that BirdFlow models produced by the new tuning framework achieve biologically realistic performance, even for prediction horizons of thousands of kilometers and several months. When species-specific data are unavailable, models can still be tuned using data from other phylogenetically adjacent species to achieve improved performance. Conclusions By integrating eBird Status & Trends abundance surfaces with data from banding recaptures, radio telemetry, and GPS tracking, we scale BirdFlow model to 153 North American migratory species, representing the first collection of continental-scale population-level movement and forecasting models. Species-specific tuning improves population-level movement forecasts, while taxonomically informed hyperparameter transfer supports the modeling of data-limited species. Overall, our work offers a foundation for more accurate predictions across hundreds of species for research in ecology and conservation, disease surveillance, aviation, and public outreach.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".