Illuminating the microbiome’s dark matter: a functional genomic toolkit for the study of human gut Actinobacteria
Bibliographic record
Abstract
Despite the remarkable evolutionary and metabolic diversity found within the human microbiome, the vast majority of mechanistic studies focus on two phyla: the Bacteroidetes and the Proteobacteria. Generalizable tools for studying the other phyla are urgently needed in order to transition microbiome research from a descriptive to a mechanistic discipline. Here, we focus on the Coriobacteriia class within the Actinobacteria phylum, detected in the distal gut of 90% of adult individuals around the world, which have been associated with both chronic and infectious disease, and play a key role in the metabolism of pharmaceutical, dietary, and endogenous compounds. We established, sequenced, and annotated a strain collection spanning 14 genera, 8 decades, and 3 continents, with a focus on Eggerthella lenta . Genome-wide alignments revealed inconsistencies in the taxonomy of the Coriobacteriia for which amendments have been proposed. Re-sequencing of the E. lenta type strain from multiple culture collections and our laboratory stock allowed us to identify errors in the finished genome and to identify point mutations associated with antibiotic resistance. Analysis of 24 E. lenta genomes revealed an “open” pan-genome suggesting we still have not fully sampled the genetic and metabolic diversity within this bacterial species. Consistent with the requirement for arginine during in vitro growth, the core E. lenta genome included the arginine dihydrolase pathway. Surprisingly, glycolysis and the citric acid cycle was also conserved in E. lenta despite the lack of evidence for carbohydrate utilization. We identified a species-specific marker gene and validated a multiplexed quantitative PCR assay for simultaneous detection of E. lenta and specific genes of interest from stool samples. Finally, we demonstrated the utility of comparative genomics for linking variable genes to strain-specific phenotypes, including antibiotic resistance and drug metabolism. To facilitate the continued functional genomic analysis of the Coriobacteriia, we have deposited the full collection of strains in DSMZ and have written a general software tool (ElenMatchR) that can be readily applied to novel phenotypic traits of interest. Together, these tools provide a first step towards a molecular understanding of the many neglected but clinically-relevant members of the human gut microbiome.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".