A call for the preservation of images, recordings, and other data in association with avian genetic samples, and the introduction of a solution: OMBIRDS
Bibliographic record
Abstract
Much current and historical research in ornithology employs catch-and-release methods, resulting in a variety of data and materials from birds for which whole-body specimens have not been collected. Often, a genetic specimen (e.g., blood or feathers) is collected along with “media specimens” such as images and/or sound recordings, providing a rich source of research material as well as an opportunity to use each type of specimen as a source of validation of the other. Despite the abundance of these datasets and their potential use in future research, the preservation of such data and associated materials is currently a task that each researcher must confront individually, which results in the loss of these research materials over time. To promote the long-term utility of information collected from the thousands of birds that are captured and released each year, we present a protocol and database template (OMBIRDS; the Online Museum of Bird Images, Recordings, and DNA Samples) for organizing and preserving images, recordings, and data associated with genetic samples. This protocol can be used by individual researchers and institutions to organize their own collections, and it also facilitates submission of records to international data repositories such as VertNet. By contributing OMBIRDS to the research community as a free database tool that can be downloaded and adapted by researchers and institutions, we hope to encourage the collection of media along with genetic samples and to facilitate the archiving of these materials for their use in future research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".