Identification and classification of the genus Bacteroides by multilocus sequence analysis
Bibliographic record
Abstract
Multilocus sequence analysis (MLSA) was performed on representative species of the genus Bacteroides. Internal fragments of the genes selected, dnaJ, gyrB, hsp60, recA, rpoB and 16S rRNA, were amplified by direct PCR and then sequenced from 38 Bacteroides strains representing 35 species. Neighbour-joining (NJ), maximum-likelihood (ML) and maximum-parsimony (MP) phylogenies of the individual genes were compared. The data confirm that the potential for discrimination of Bacteroides species is greater using MLSA of housekeeping genes than 16S rRNA genes. Among the housekeeping genes analysed, gyrB was the most informative, followed by dnaJ. Analyses of concatenated sequences (4816 bp) of all six genes revealed robust phylogenetic relationships among different Bacteroides species when compared with the single-gene trees. The NJ, ML and MP trees were very similar, and almost fully resolved relationships of Bacteroides species were obtained, to our knowledge for the first time. In addition, analysis of a concatenation (2457 bp) of the dnaJ, gyrB and hsp60 genes produced essentially the same result. Ten distinct clades were recognized using the SplitsTree4 program. For the genus Bacteroides, we can define species as a group of strains that share at least 97.5% gene sequence similarity based on the fragments of five protein-coding housekeeping genes and the 16S rRNA gene. This study demonstrates that MLSA of housekeeping genes is a valuable alternative technique for the identification and classification of species of the genus Bacteroides.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".