MétaCan
Menu
Back to cohort
Record W4411691518 · doi:10.1093/gigascience/giaf072

Integrating comparative genomics and risk classification by assessing virulence, antimicrobial resistance, and plasmid spread in microbial communities with gSpreadComp

2025· article· en· W4411691518 on OpenAlexfundno aff
Jonas Coelho Kasmanas, Stefanía Magnúsdóttir, Junya Zhang, Kornelia Smalla, Michael Schloter, Peter F. Stadler, André C. P. L. F. de Carvalho, Ulisses Nunes da Rocha

Bibliographic record

VenueGigaScience · 2025
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGut microbiota and health
Canadian institutionsnot available
FundersGerman Network for Bioinformatics InfrastructureFundação de Amparo à Pesquisa do Estado de São PauloDeutsche ForschungsgemeinschaftInternational Development Research CentreAlexander von Humboldt-Stiftung
KeywordsMetagenomicsBiologyVirulenceGenomicsComparative genomicsGeneticsAntibiotic resistanceWorkflowGenomeHorizontal gene transferComputational biologyGeneComputer scienceDatabase

Abstract

fetched live from OpenAlex

BACKGROUND: Comparative genomics, genetic spread analysis, and context-aware ranking are crucial in understanding microbial dynamics' impact on public health. gSpreadComp streamlines the path from in silico analysis to hypothesis generation. By integrating comparative genomics, genome annotation, normalization, plasmid-mediated gene transfer, and microbial resistance-virulence risk-ranking into a unified workflow, gSpreadComp facilitates hypothesis generation from complex microbial datasets. FINDINGS: The gSpreadComp workflow works through 6 modular steps: taxonomy assignment, genome quality estimation, antimicrobial resistance (AMR) gene annotation, plasmid/chromosome classification, virulence factor annotation, and downstream analysis. Our workflow calculates gene spread using normalized weighted average prevalence and ranks potential resistance-virulence risk by integrating microbial resistance, virulence, and plasmid transmissibility data and producing an HTML report. As a use case, we analyzed 3,566 metagenome-assembled genomes recovered from human gut microbiomes across diets. Our findings indicated consistent AMR across diets, with diet-specific resistance patterns, such as increased bacitracin in vegans and tetracycline in omnivores. Notably, ketogenic diets showed a slightly higher resistance-virulence rank, while vegan and vegetarian diets encompassed more plasmid-mediated gene transfer. CONCLUSIONS: The gSpreadComp workflow aims to facilitate hypothesis generation for targeted experimental validations by the identification of concerning resistant hotspots in complex microbial datasets. Our study raises attention to a more thorough study of the critical role of diet in microbial community dynamics and the spread of AMR. This research underscores the importance of integrating genomic data into public health strategies to combat AMR. The gSpreadComp workflow is available at https://github.com/mdsufz/gSpreadComp/.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.010
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.005
Threshold uncertainty score0.027

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.010
Meta-epidemiology (narrow)0.0030.001
Meta-epidemiology (broad)0.0010.004
Bibliometrics0.0050.003
Science and technology studies0.0010.001
Scholarly communication0.0040.002
Open science0.0020.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0050.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.015
GPT teacher head0.280
Teacher spread0.265 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueGigaScienceSame topicGut microbiota and healthFrench-language works237,207