pEX Codebase: Anthropocene Imperilment of Ancient Diversity and Evolutionary Potential in Terrestrial Vertebrates
Notice bibliographique
Résumé
Zenodo Readme – Pyron et al. pEX Codebase This repository contains the complete codebase to recreate the analyses in the manuscript. It is based on Zenodo repositories for the TetrapodTraits attribute dataset and the TetrapodTrees phylogenetic dataset, from which we used 100 randomly sampled trees. Folders (provided as *.zip files, preserving internal directory structures): Attributes: Contains copies of the attribute data from the TetrapodTraits dataset, including the new data and traits introduced in this MS. Files in this folder are primarily utilized in Step 2. Note: This archive contains the versions of the TetrapodTraits attribute dataset and the sample of 100 trees from the TetrapodTrees phylogenetic dataset used in the published version of the pEX models for reproducibility. If you want to create new models or run other types of analyses, always check the main repositories for those datasets to obtain the latest version. Features: Contains the output from the hyperparameter optimization and feature selection in Step 3. Figures: Contains the figures generated from the *.Rmd document representing the final analyses. Graphics: Contains the phylogeny graphics generated in Step 6 Maps: Contains the eigenvalues, grid cells, and randomized assemblages used to generate the maps, which are also printed to this directory. Models: Contains the pEX and BRMS models from Steps 4 and 5. Output: Contains all the final output metrics generated by Step 7. PCA: Contains the first 100 PC axes of 100 randomly sampled trees from the TetrapodTrees database. Predictors: Contains the attribute files created by combining the static and imputed TetrapodTraits trait data with spatial filters and phylogenetic PC axes used in the pEX model training in xgboost, generated in Step 2. Trees: Contains 100 randomly sampled trees from the TetrapodTrees database used in Step 1 to generate the files in the PCA directory. Files: tetrapoda_1.0_pEX_summary.html: HTML markdown document outlining the workflow and summarizing the key results, along with figures. tetrapoda_1.0_pEX_summary.Rmd: R markdown code collecting the final outputs, figures, and summaries, and providing final statistical analyses of various metrics. tetrapoda_1.0_step1_PCA.R: Calculates 100 PC axes from the phylogenetic variance-covariance matrices of 100 randomly sampled trees from the TetrapodTrees database in the Trees directory, saved to the PCA directory. tetrapoda_1.0_step2_predictors.R: Takes the various attributes and spatial filters in the Attributes directory and PC axes in the PCA directory, and combines them to produce input files for model training in the Predictors directory. tetrapoda_1.0_step3_features.R: Optimizes hyperparameters and performs feature selection on the combined attributes, saved to the Features directory. tetrapoda_1.0_step4_pEX.R: Optimizes 100 pEX models in xgboost using the input files from the Predictors directory, saved to the Models directory and summarized in the Outputs directory. tetrapoda_1.0_step5_brms.R: Fits 196 beta regression models and 25 clade-specific models linking pEX to ED, DR, and Clade using median values saved in the Outputs directory. tetrapoda_1.0_step6_plot.R: Plots the median values for pEX, ED, DR, and threat status on a randomly sampled tree from the TetrapodTrees database in the Trees directory, saved to the Graphics directory. tetrapoda_1.0_step7_outputs.R: Summarizes and combines all metrics (pEX, ED, DR, and threat status) into the Outputs directory, at the species and family level, along with variable importance. tetrapoda_1.0_step8_maps.R: Generates the maps in the MS and Extended Data based on median pEX, ED, DR, and range rarity, including randomized assemblages to account for species richness, and across latitudes. Part of the VertLife initiative: An NSF-sponsored, multi-institutional project to study the biodiversity of all terrestrial vertebrates (Tetrapoda), making them the first major global group of animals with near-complete species-level data on key evolutionary and ecological attributes. Contact: Alex Pyron (rpyron@gwu.edu), Mario Moura (mariormoura@gmail.com), and Walter Jetz (wjetz@yale.edu). We thank the Map of Life team at the Yale Center for Biodiversity and Global Change for their contribution to earlier versions of this dataset and the E.O. Wilson Biodiversity Foundation for support in furtherance of the Half-Earth Project to W.J. This work was supported by funding from US National Science Foundation (NSF) grants to R.A.P. (DBI-0905765, DEB-1441719), R.C.K.B. (DEB-1441652), T.J.C. (DBI-2334779, DEB-2406685), J.E. (DEB-1441634, DEB-2244754), R.P.G (DEB-1441628), and W.J. (DEB-1441737). Additional funding came from US National Institutes of Health (NIH) grants to N.S.U. (1R21AI164268 and 1R35GM156919); US National Aeronautics and Space Administration (NASA) grants to W.J. (80NSSC17K0282 and 80NSSC18K0435); BR São Paulo Research Foundation (FAPESP) grants to M.R.M. (#2021/11840-6 and #2022/12231-6), K.C. (#2020/12558-0), and M.T.M. (#2023/14506-5); BR Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES) fellowships for J.J.M.G.; BR Fonseca Leadership Program (GEF/FUNBIO #108/2025) to J.P.O.X., BR Conselho Nacional de Desenvolvimento Científico (CNPq) grants to K.C. (#444240/2024-1); NSERC Canada grants to A.Ø.M.; and a UCLA Chancellor’s Fellowship to R.M.P.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,011 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,005 | 0,005 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,004 | 0,003 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,168 | 0,094 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».