Taxonomy of <i>Rhizobiaceae</i> revisited: proposal of a new framework for genus delimitation
Bibliographic record
Abstract
ABSTRACT The alphaproteobacterial family Rhizobiaceae is highly diverse, with 168 species with validly published names classified into 17 genera with validly published names. Most named genera in this family are delineated based on genomic relatedness and phylogenetic relationships, but some historically named genera show inconsistent distribution and phylogenetic breadth. Most problematic is Rhizobium , which is notorious for being highly paraphyletic, as most newly described species in the family being assigned to this genus without consideration for their proximity to existing genera, or the need to create novel genera. In addition, many Rhizobiaceae genera lack synapomorphic traits that would give them biological and ecological significance. We propose a common framework for genus delimitation within the family Rhizobiaceae . We propose that genera in this family should be defined as monophyletic groups in a core-genome gene phylogeny, that are separated from related species using a pairwise core-proteome average amino acid identity (cpAAI) threshold of approximately 86%. We further propose that the presence of additional genomic or phenotypic evidence can justify the division of species into separate genera even if they all share greater than 86% cpAAI. Applying this framework, we propose to reclassify Rhizobium rhizosphaerae and Rhizobium oryzae into the new genus Xaviernesmea gen. nov. Data is also provided to support the recently proposed genus “ Peteryoungia ”, and the reclassifications of Rhizobium yantingense as Endobacterium yantingense comb. nov., Rhizobium petrolearium as Neorhizobium petrolearium comb. nov., Rhizobium arenae as Pararhizobium arenae comb. nov., Rhizobium tarimense as Pseudorhizobium tarimense comb. nov., and Rhizobium azooxidefex as Mycoplana azooxidifex comb. nov. Lastly, we present arguments that the unification of the genera Ensifer and Sinorhizobium in Opinion 84 of the Judicial Commission is no longer justified by current genomic and phenotypic data. We thus argue that the genus Sinorhizobium is not illegitimate and now encompasses 17 species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.002 |
| Science and technology studies | 0.002 | 0.004 |
| Scholarly communication | 0.004 | 0.005 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".