The phonology of sperm whale coda vowels
Bibliographic record
Abstract
ABSTRACT In previous research, sperm whale codas (structured series of clicks used for communication) have been shown to resemble human vowels acoustically. Based on the number of formants, two different coda quality categories have been described: a -codas and i -codas. In the present paper, we demonstrate that sperm whale codas not only resemble human vowels acoustically, but also pattern like them across several dimensions. First, traditional count- and timing-based coda types interact with coda “vowel” quality ( a vs. i ). Second, a -codas are generally longer than i -codas. Third, the duration of i -codas has a bimodal distribution, showing a contrast between short i -codas and long ī -codas. Fourth, the baseline coda length differs across whales. And fifth, edge clicks mismatching their coda often match an adjacent coda, a phenomenon that resembles human coarticulation. All five properties have close parallels in the phonetics and phonology of human languages. Sperm whale coda vocalizations thus represent one of the closest parallels to human phonology of any known animal communication system. SIGNIFICANCE STATEMENT Sperm whales communicate using series of clicks known as codas . The codas acoustically resemble human vowels. In addition, they pattern in ways similar to human sound systems. For example, different coda types are correlated with particular click qualities, and their durations are intentionally controlled. This shows that sperm whale vocalizations are highly complex and likely constitute one of the most sophisticated communication systems in the animal kingdom. By studying it, we may be able to gain a broader understanding of animal intelligence and social behaviors, determine the impact of human activities on whale habitat, develop strategies to protect whales from threats such as noise pollution and ship traffic, and advance legislation which facilitates that protection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".