Identification of a botulinum neurotoxin-like gene cluster in <i>Bacillus toyonensis</i>
Bibliographic record
Abstract
Abstract Clostridial neurotoxins, which include botulinum neurotoxins (BoNTs) and tetanus neurotoxin (TeNT) are the most potent toxins known, and are the causative agents of the neuroparalytic diseases, botulism and tetanus. Until recently, the clostridial neurotoxin family was restricted to the genus Clostridium , but members of this protein family have been found in a growing number of non- Clostridium species including Weissella, Enterococcus , and Paraclostridium . Here, we report the bioinformatic identification and analysis of a novel clostridial neurotoxin homolog in a Bacillus toyonensis genome recently deposited into the NCBI Genbank database. This putative toxin shares 26-29% identity with its closest BoNT relatives, suggesting that it is likely a novel BoNT-like toxin. It possesses key functional motifs (e.g., HExxH) indicative of toxin protease activity, contains the four characteristic BoNT domains, and is located in a BoNT-like genomic neighborhood containing the upstream non-toxic non-hemagglutinin (NTNH) gene as well as several P47-related genes. Phylogenetically, the toxin clusters as a divergent member of the recently discovered lineage of BoNT-like toxins that includes BoNT/X, BoNT/En, and the insecticidal PMP1. Genomic analysis of the B. toyonensis isolate CH177 revealed additional virulence factors and toxin genes indicative of potential pathogenicity targeting an unknown host species. The Bacillus toyonensis BoNT-like protein (BTNT) adds to a growing number of non-clostridial BoNT-like toxins, adding further information on the intriguing phylogenetic distribution and evolutionary history of the most potent toxins known.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".