OTUs clustering should be avoided for defining oral microbiome
Bibliographic record
Abstract
ABSTRACT This in silico investigation aimed to: 1) evaluate a set of primer pairs with high coverage, including those most commonly used in the literature, to find the different oral species with 16S rRNA gene amplicon similarity/identity (ASI) values ≥97%; and 2) identify oral species that may be erroneously clustered in the same operational taxonomic unit (OTU) and ascertain whether they belong to distinct genera or other higher taxonomic ranks. Thirty-nine primer pairs were employed to obtain amplicon sequence variants (ASVs) from the complete genomes of 186 bacterial and 135 archaeal species. For each primer, ASVs without mismatches were aligned using BLASTN and their similarity values were obtained. Finally, we selected ASVs from different species with an ASI value ≥97% that were covered 100% by the query sequences. For each primer, the percentage of species-level coverage with no ASI≥97% (SC-NASI≥97%) was calculated. Based on the SC-NASI≥97% values, the best primer pairs were OP_F053-KP_R020 for bacteria (65.05%), KP_F018-KP_R002 for archaea (51.11%), and OP_F114-KP_R031 for bacteria and archaea together (52.02%). Eighty percent of the oral-bacteria and oralarchaea species shared an ASI≥97% with at least one other taxa, including Campylobacter , Rothia , Streptococcus , and Tannerella , which played conflicting roles in the oral microbiota. Moreover, around a quarter and a third of these two-by-two similarity relationships were between species from different bacteria and archaea genera, respectively. Furthermore, even taxa from distinct families, orders, and classes could be grouped in the same cluster. Consequently, irrespective of the primer pair used, OTUs constructed with a 97% similarity provide an inaccurate description of oral-bacterial and oral-archaeal species, greatly affecting microbial diversity parameters. As a result, clustering by OTUs impacts the credibility of the associations between some oral species and certain health and disease conditions. This limits significantly the comparability of the microbial diversity findings reported in oral microbiome literature.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".