Hybrid sequence-based analysis reveals the distribution of bacterial species and genes in the oral microbiome at a high resolution
Bibliographic record
Abstract
Bacteria in the oral microbiome are poorly identified owing to the lack of established culture methods for them. Thus, this study aimed to use culture-free analysis techniques, including bacterial single-cell genome sequencing , to identify bacterial species and investigate gene distribution in saliva. Saliva samples from the same individual were classified as inactivated or viable and then analyzed using 16S rRNA sequencing, metagenomic shotgun sequencing , and bacterial single-cell sequencing. The results of 16S rRNA sequencing revealed similar microbiota structures in both samples, with Streptococcus being the predominant genus. Metagenomic shotgun sequencing showed that approximately 80 % of the DNA in the samples was of non-bacterial origin, whereas single-cell sequencing showed an average contamination rate of 10.4 % per genome. Single-cell sequencing also yielded genome sequences for 43 out of 48 wells for the inactivated samples and 45 out of 48 wells for the viable samples. With respect to resistance genes, four out of 88 isolates carried cfxA , which encodes a β-lactamase, and four isolates carried erythromycin resistance genes. Tetracycline resistance genes were found in nine bacteria. Metagenomic shotgun sequencing provided complete sequences of cfxA , ermF , and ermX , whereas other resistance genes, such as tetQ and tetM , were detected as fragments. In addition, virulence factors from Streptococcus pneumoniae were the most common, with 13 genes detected. Our average nucleotide identity analysis also suggested five single-cell-isolated bacteria as potential novel species. These data would contribute to expanding the oral microbiome data resource.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".