skDER & CiDDER: two scalable approaches for microbial genome dereplication
Notice bibliographique
Résumé
ABSTRACT An abundance of microbial genomes have been sequenced in the past two decades. For fundamental comparative genomic investigations, where the goal is to determine the major gain and loss events shaping the pangenome of a species, it is often unnecessary and computationally onerous to include all available genomes in studies. In addition, over-representation of specific lineages due to sampling and sequencing bias can have undesired effects on evolutionary analyses. To assist users with genomic dereplication , selecting a subset of representative genomes, for downstream comparative genomic investigations, we developed skDER & CiDDER ( https://github.com/raufs/skDER ). skDER combines recent advances to efficiently estimate average nucleotide identity (ANI) between thousands of microbial genomes with two efficient algorithms for genomic dereplication. Further, CiDDER implements an approach whereby protein clusters are determined across all genomes and genomes are iteratively selected as representatives until a user-defined saturation of the total protein space is achieved. To support ease of use, several auxiliary functionalities are implemented within the two programs, including arguments to: (i) test the number of representative genomes resulting from a variety of clustering parameters, (ii) automate downloading of genomes belonging to a bacterial species or genus by name, (iii) cluster non-representative genomes to their closest representative genomes, and (iv) automatically filter predicted plasmids and phages prior to dereplication. We further assess the effects of filtering mobile genetic elements (MGEs) on ANI and alignment fraction (AF) estimates between pairs of genomes and find that MGEs tend to slightly deflate both metrics in one species. DATA SUMMARY skDER and CiDDER are provided as open-source software implemented in Python and C++ on Github: https://github.com/raufs/skDER ; with version updates tracked on Zenodo: https://zenodo.org/records/13887710 1 . Installation of the software is supported via both Bioconda 2 and Docker. Additional code and data for analyses presented in this manuscript can be found on Zenodo at: https://zenodo.org/records/13891800 3 . Pre-computed representative genomes selected by skDER (v1.0.7) for 18 common bacterial taxonomic groups, referencing classifications from GTDB release 214 4 , are also provided on Zenodo at: https://zenodo.org/records/10041203 5 . IMPACT STATEMENT Due to the increased availability of genomes for certain microbial species, performing fundamental comparative genomic investigations has become restricted to those with access to more advanced computational infrastructure. Genomic dereplication, the process of selecting distinct representative genomes to capture the breadth of a taxonomic group, thus presents a valuable solution to overcome this embarrassment of riches. Specifically, genomic dereplication allows simplifying the scale of comparative investigations while minimizing the risk of biasing analyses to specific lineages, which might be overrepresented in genomic databases. We present here two programs for genomic dereplication, one based on ANI-inference, skDER, and the other based on assessing the saturation of total coding-genes sampled, CIDDER. These tools are implemented with a variety of auxiliary options and designed for ease of use.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,013 |
| Méta-épidémiologie (sens strict) | 0,003 | 0,002 |
| Méta-épidémiologie (sens large) | 0,002 | 0,003 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,002 | 0,001 |
| Communication savante | 0,003 | 0,004 |
| Science ouverte | 0,006 | 0,008 |
| Intégrité de la recherche | 0,002 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,013 | 0,012 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».