Code and data for: Traits, threats, and popularity explain extinction risk of bird
Notice bibliographique
Résumé
README for Supplementary MaterialsTitle of Paper: Traits, threats, and popularity explain extinction risk of birds globally Authors: Janaina Serrano, Lars Iversen, Laura Pollock DescriptionThis README file provides an overview of the dataset, columns, and code files included in the supplementary materials for the paper titled Traits, threats, and popularity explain extinction risk of birds globally. The dataset contains information on bird species, their traits, threats they face, habitat characteristics, and popularity metrics. The associated code files are provided to facilitate reproducibility and further exploration.Files in Supplementary Materials1. Dataset FileFilename: datFormat: CSVDataset DescriptionColumns in the DatasetColumn NameDescriptionspeciesScientific name of the bird species.threatenedBinary indicator of whether the species is classified as threatened (1/0).HabitatPrimary habitat type for the species.Order1Taxonomic order to which the species belongs.total_threatsTotal number of threats affecting the species.MigrationMigration status of the species (e.g., migratory or non-migratory).HabitatBreadthBreadth of habitat types the species can occupy (numerical value).Range.SizeGeographical range size of the species (in square kilometers).MassBody mass of the species (in grams).gbif_meanMean number of GBIF records indicating species' data availability.meanMean number of Google search hits for the species (popularity metric).pollutionImpact of pollution threats on the species.loggingImpact of logging threats on the species.invasiveImpact of invasive species threats on the species.agricultureImpact of agricultural activities on the species.climate_changeImpact of climate change threats on the species.huntingImpact of hunting threats on the species.Mass.logLog-transformed body mass.Range.Size.logLog-transformed range size.gbif_mean.logLog-transformed GBIF mean value.HabitatBreadth.logLog-transformed habitat breadth.google.logLog-transformed Google search popularity.Migration1Categorical indicator of migration type.Mass.log.csCentered and scaled log-transformed body mass.Range.Size.log.csCentered and scaled log-transformed range size.gbif_mean.log.csCentered and scaled log-transformed GBIF mean value.HabitatBreadth.log.csCentered and scaled log-transformed habitat breadth.google.log.csCentered and scaled log-transformed Google popularity. 2. Code FilesDescriptions of the provided code files:FilenameDescriptiondataprep_birdtraitsScript for preparing and cleaning the bird trait dataset. Includes processes like data wrangling, standardization of column names, log-transformations of variables, and preparation of final input data for analysis.model_figures_scriptScript to generate the main figures from the paper. This includes plotting extinction risk models, trait relationships, and visualizations of the species' threats and popularity metrics.model_evaluation_bivmapScript for model evaluation and bivariate mapping. Evaluates the predictive performance of models and creates spatial maps to visualize overlaps modeled extinction risk and observed conservation status of birds.gtrendsScript for collecting and processing Google Trends data related to species' popularity. Retrieves search hit data, processes it for analysis, and calculates summary metrics like mean hits and log-transformed values.GBIF_download Script for collecting and processing GBIF number of observations for birds globally from 2004-2021. Retrieves GBIF data, processes it for analysis, and calculates average species observations per year.Data Usage and CitationThe dataset is provided as supplementary material for the paper and can be used for academic purposes. If you use this dataset, please cite the paper as follows:Serrano J., Iversen L., Pollock L. (2025). Traits, threats, and popularity explain extinction risk of birds globally.Contact InformationFor questions about the dataset or paper, please contact:Janaina Serrano: janaina.serrano@mail.mcgill.caLars Iversen: lars.iversen@mcgill.caLaura Pollock: laura.pollock@mcgill.ca
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,013 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,002 | 0,002 |
| Science ouverte | 0,002 | 0,002 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,335 | 0,222 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».