A Global Library of Underwater Biological Sounds (GLUBS): An Online Platform with Multiple Passive Acoustic Monitoring Applications
Bibliographic record
Abstract
Aquatic ecosystems contain some of the world’s most diverse environments, and their soundscapes are often teeming with sounds from a wealth of biological sources. Passive acoustic monitoring (PAM) is an increasingly accessible technique that offers an unprecedented, non-extractive means to “observe” these habitats, many of which are too deep, dark, turbid, or remote to sample easily with other methods. Applications to assist analysis of PAM data already exist (e.g., reference libraries, data portals, discussion forums), machine learning code is increasingly more available, and citizen science programs are broadening public interest in underwater sound. However, individually, these resources do not realize their full potential. To help address this limitation, a working group for a Global Library of Underwater Biological Sounds (GLUBS) has proposed an open-access web-based single-point-of-contact platform to integrate and expand these applications to help broaden and standardize scientific and community knowledge of underwater soundscapes and their contributing sources. This paper presents a summary of a meeting of the GLUBS working group that was held at “The Effects of Noise on Aquatic Life, 2022,” including some of the core values, initial targets, points for design, data management issues, and potential avenues for stakeholder engagement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.010 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.105 | 0.086 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".