Phylogeny of the Order Bacillales inferred from 3’ 16S rDNA and 5’ 16S-23S ITS nucleotide sequences
Bibliographic record
Abstract
A short 220 bp sequence was used to study the taxonomic organization of the bacterial Order Bacillales. The nucleotide sequences of the 3' end of the 16S rDNA and the 16S-23S Internal transcribed spacer (ITS) were determined for 32 Bacillales species and strains. The data for 40 additional Bacillales species and strains were retrieved directly from Genbank. Together, these 72 Bacillales species and strains encompassed eight families and 21 genera. The 220 bp sequence used here covers a conserved 150 bp sequence located at the 3' end of the 16S rDNA and a conserved 70 bp sequence located at the 5' end of the 16S-23S ITS. A neighbor-joining phylogenetic tree was inferred from comparative analyses of all 72 nucleotide sequences. Eight major Groups were revealed. Each Group was sub-divided into sub-groups and branches. In general, the neighbor-joining tree presented here is in agreement with the currently accepted phylogeny of the Order Bacillales based on phenotypic and genotypic data. The use of this 220 bp sequence for phylogenetic analyses presents several advantages over the use of the entire 16S rRNA genes or the generation of extensive phenotypic and genotypic data. This 220 bp sequence contains 150 bp at the 3' end of the 16S rDNA which allows discrimination among distantly related species and 70 bp at the 5' end of the 16S-23S ITS which, owing to its higher percentage of nucleotide sequence divergence, adds discriminating power among closely related species from same genus and closely related genera from same family. The method is simple, rapid, suited to large screening programs and easily accessible to most laboratories.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".