Ribotype Classification of Clostridioides difficile Isolates Is Not Predictive of the Amino Acid Sequence Diversity of the Toxin Virulence Factors TcdA and TcdB
Bibliographic record
Abstract
Clostridioides (Clostridium) difficile is the most commonly recognized cause of infectious diarrhea in healthcare settings. Currently there is no vaccine to prevent initial or recurrent C. difficile infection (CDI). Two large clostridial toxins, TcdA and TcdB, are the primary virulence factors for CDI. Immunological approaches to prevent CDI include antibody-mediated neutralization of the cytotoxicity of these toxins. An understanding of the sequence diversity of the two toxins expressed by disease causing isolates is critical for the interpretation of the immune response to the toxins. In this study, we determined the whole genome sequence (WGS) of 478 C. difficile isolates collected in 12 countries between 2004-2018 to probe toxin variant diversity. A total of 44 unique TcdA variants and 37 unique TcdB variants were identified. The amino acid sequence conservation among the TcdA variants (> 98%) is considerably greater than among the TcdB variants (as low as 86.1%), suggesting that different selection pressures may have contributed to the evolution of the two toxins. Phylogenomic analysis of the WGS data demonstrate that isolates grouped together based on ribotype or MLST code for multiple different toxin variants. These findings illustrate the importance of determining not only the ribotype but also the toxin sequence when evaluating strain coverage using vaccine strategies that target these virulence factors. We recommend that toxin variant type and sequence type (ST), be used together with ribotype data to provide a more comprehensive strain classification scheme for C. difficile surveillance during vaccine development objectives.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".