AvianLexiconAtlas: A database of descriptive categories of English-language bird names around the world
Notice bibliographique
Résumé
Abstract Common names of species are important for communicating with the general public. In principle, these names should provide an accessible way to engage with and identify species. The common names of species have historically been labile without standard guidelines, even within a language. Currently, there is no systematic assessment of how often common names communicate identifiable and biologically relevant characteristics about species. This is a salient issue in ornithology, where common names are used more often than scientific names for species of birds in written and spoken English, even by professional researchers. To gain a better understanding of the types of terminology used in the English-language common names of bird species, a group of 85 professional ornithologists and non-professional contributors classified unique descriptors in the common names of all recognized species of birds. In the AvianLexiconAtlas database produced by this work, each species’ common name is assigned to one of ten categories associated with aspects of avian biology, ecology, or human culture. Across 10,906 species of birds, 89% have names describing the biology of the species, while the remaining 11% of species have names derived from human cultural references, human names, or local non-English languages. Species with common names based on features of avian biology are more likely to be related to each other or be from the same geographic region. The crowdsourced data collection also revealed that many common names contain specialized or historic terminology unknown to many of the data collectors, and we include these terms in a glossary and gazetteer alongside the dataset. The AvianLexiconAtlas can be used as a quantitative resource to assess the state of terminology in English-language common names of birds. Future research using the database can shed light on historical approaches to nomenclature and how people engage with species through their names.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».