Minimum Core Genome Sequence Typing of Bacterial Pathogens: a Unified Approach for Clinical and Public Health Microbiology
Bibliographic record
Abstract
Bacterial pathogens impose a heavy health burden worldwide. In the new era of high-throughput sequencing and online bioinformatics, real-time genome typing of infecting agents, and in particular those with potential severe clinical outcomes, holds promise for guiding clinical care to limit the detrimental effects of infections and to prevent potential local or global outbreaks. Here, we sequenced and compared 85 isolates of Streptococcus suis, a zoonotic human and swine pathogen, wherein we analyzed 32 recognized serotypes and 75 sequence types representing the diversity of the species and the human clinical isolates with high public health significance. We found that 1,077 of the 2,469 genes are shared by all isolates. Excluding 201 common but mobile genes, 876 genes were defined as the minimum core genome (MCG) of the species. Of 190,894 single-nucleotide polymorphisms (SNPs) identified, 58,501 were located in the MCG genes and were referred to as MCG SNPs. A population structure analysis of these MCG SNPs classified the 85 isolates into seven MCG groups, of which MCG group 1 includes all isolates from human infections and outbreaks. Our MCG typing system for S. suis provided a clear separation of groups containing human-associated isolates from those containing animal-associated isolates. It also separated the group containing outbreak isolates, including those causing life-threatening streptococcal toxic shock-like syndrome, from sporadic or less severe meningitis or bacteremia-only isolates. The typing system facilitates the application of genome data to the fields of clinical medicine and epidemiology and to the surveillance of S. suis. The MCG groups may also be used as the taxonomical units of S. suis to define bacterial subpopulations with the potential to cause severe clinical infections and large-scale outbreaks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".