Integrated population clustering and genomic epidemiology with PopPIPE
Notice bibliographique
Résumé
Abstract Genetic distances between bacterial DNA sequences can be used to cluster populations into closely related subpopulations, and as an additional source of information when detecting possible transmission events. Due to their variable gene content and order, reference-free methods offer more sensitive detection of genetic differences, especially among closely related samples found in outbreaks. However, across longer genetic distances, frequent recombination can make calculation and interpretation of these differences more challenging, requiring significant bioinformatic expertise and manual intervention during the analysis process. Here we present a Pop ulation analysis PIPE line (PopPIPE) which combines rapid reference-free genome analysis methods to analyse bacterial genomes across these two scales, splitting whole populations into subclusters and detecting plausible transmission events within closely related clusters. We use k-mer sketching to split populations into strains, followed by split k-mer analysis and recombination removal to create alignments and subclusters within these strains. We first show that this approach creates high quality subclusters on a population-wide dataset of Streptococcus pneumoniae . When applied to nosocomial vancomycin resistant Enterococcus faecium samples, PopPIPE finds transmission clusters which are more epidemiologically plausible than core genome or MLST-based approaches. Our pipeline is rapid and reproducible, creates interactive visualisations, and can easily be reconfigured and re-run on new datasets. Therefore PopPIPE provides a user-friendly pipeline for analyses spanning species-wide clustering to outbreak investigations. Impact statement As time passes, bacterial genomes accumulate small changes in their sequence due to mutations, or larger changes in their content due to horizontal gene transfer. Using their genome sequences, it is possible to use phylogenetics to work out the most likely order in which these changes happened, and how long they took to happen. Then, one can estimate the time that separates any two bacterial samples – if it is short then they may have been directly transmitted or acquired from the same source; but if it is long they must have been acquired separately. This information can be used to determine transmission chains, in conjunction with dates and locations of infections. Understanding transmission chains enables targeted infection control measures. However, correctly calculating the genetic evidence for transmission is made difficult by correctly distinguishing different types of sequence changes, dealing with large amounts of genome data, and the need to use multiple complex bioinformatic tools. We addressed this gap by creating a computational workflow, PopPIPE, which automates the process of detecting possible transmissions using genome sequences. PopPIPE applies state-of-the-art tools and is fast and easy to run – making this technology will be available to a wider audience of researchers. Data summary The code for this pipeline is available at https://github.com/bacpop/PopPIPE and as a docker image https://hub.docker.com/r/poppunk/poppipe . Raw sequencing reads for Enterococcus faecium isolates have been deposited at the NCBI under BioProject accession number PRJNA997588.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,011 |
| Méta-épidémiologie (sens strict) | 0,003 | 0,002 |
| Méta-épidémiologie (sens large) | 0,003 | 0,004 |
| Bibliométrie | 0,003 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,003 | 0,002 |
| Science ouverte | 0,004 | 0,004 |
| Intégrité de la recherche | 0,001 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,017 | 0,006 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».