Best of Both Worlds? Optimising Graph-Based Antimicrobial Resistance Gene Profiling in Long and Short-Read Metagenomes
Notice bibliographique
Résumé
Abstract Environmental surveillance using metagenomic sequencing offers a powerful way to track emerging and mobile antimicrobial resistance (AMR) genes and inform public health mitigation strategies. Read-based analysis tools can sensitively detect AMR genes in metagenomes but provide little information about the surrounding genome. This prevents easily linking detected genes with particular host species or mobile genetic elements. On the other hand, contig-based analysis tools can provide this genomic context but systematically fail to recover many AMR genes. Querying the intermediate assembly graph directly may provide a trade-off between these strengths and weaknesses. However, many existing tools capable of querying assembly graphs are designed for applications other than gene detection, such as pan-genomics, indexing, or scaffolding. Therefore, a comprehensive evaluation of six tools across four search paradigms was performed to determine the optimal graph querying tool for profiling AMR genes in both long and short-read metagenomic assembly graphs. Across mock and simulated metagenomes of varying complexity and read-type, BLAST-based graph alignment (as implemented by GraphAligner) consistently outperformed other graph alignment algorithms. Overall, graph-based methods correctly identified 21% to 46% more AMR genes in complex datasets than contig analyses; however, increases in recall were modest. Combining assembly graphs analyses with contig-based analyses identifies up to 56% additional AMR genes across both long and short-read datasets. This study highlights the challenges associated with metagenomic AMR surveillance and demonstrates that graph-based analyses offer a useful tool in maximising sensitive identification of AMR genes and their genomic context from these data. Importance Antimicrobial resistance (AMR) is a severe public health threat that has spurred non-governmental organisations and public health agencies to develop action plans to reduce resistance to critical antimicrobials. Surveillance of One Health environments for AMR determinants are often central parts of these action plans. Metagenomic sequencing presents a key method for clinical and public health AMR surveillance; however, algorithmic and biochemical limitations prevent linking most detected AMR genes to their associated host bacteria or mobile genetic elements. Our findings suggest that querying the assembly graph alongside assembled contigs can identify more AMR genes than contigs alone while still providing epidemiologically informative flanking sequences. Associating AMR genes with their genomic context greatly expands our ability to assess the risk they pose across different environments. These improvements in metagenomic AMR gene identification make AMR surveillance more effective for public health institutions potentially reducing the harm of resistant infections.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,011 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,003 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,003 | 0,004 |
| Science ouverte | 0,001 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».