Meta-analysis identifies microbial signatures of disease in murine models of inflammatory bowel disease
Bibliographic record
Abstract
ABSTRACT The gut microbiota plays a central role in modulating intestinal inflammation, but the identification of specific inflammation-associated microbes has remained elusive. Here, we perform a meta-analysis on metagenomic data from 12 different studies of murine colitis triggered by a variety of genetic and environmental factors with the goal of finding bacterial taxonomic groups that can act as signatures of health or disease across studies, and that can be used to discriminate between healthy and diseased mice. We leveraged recent developments in 16S analysis tools to identify amplicon sequence variants (ASVs) instead of the traditional Operational Taxonomic Units, and used the EZTaxon reference database that distinguishes between currently unnamed and uncharacterized 16S phylotypes. Random Forest model and differential abundance analysis were used to detect microbial signatures that could consistently differentiate healthy from diseased mice, and a ‘dysbiosis index’ was constructed from these. This dysbiosis index was able to correctly distinguish samples derived from inflamed and non-inflamed mice in the majority of studies and significantly outperformed other frequently used metrics of dysbiosis including alpha-diversity, proteobacterial abundance, and the ratio of Bacteroidetes to Firmicutes. 10 of 12 bacteria we identify as associated with the diseased state are members of the order Bacteroidales, including several species from the abundant but poorly understood S24-7 family. The implications of these findings are discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".