Strain-level and phenotypic stability contrasts with plasmid and phage variability in water kefir communities
Bibliographic record
Abstract
Abstract Microbial communities can change in response to top-down factors, such as phages, and bottom-up factors, such as nutrient availability. Previous studies have successfully investigated bacterial species-level dynamics, but diversity and interactions beyond the species-level is usually lacking. Traditional fermented foods, such as water kefir, provide ideal systems to study ecological and evolutionary dynamics beyond the species-level, as they are simple and trackable systems that are cultivated in non-sterile, nutrient-rich environments which foster microbial growth and invasion. Despite the central role of only a few lactic acid bacteria for fermentation, little is known about the genomic diversity and dynamics of these community members over time. Within the framework of a graduate course, 35 students propagated water kefir across several generations under different nutrient conditions and in different households to study microbial responses over time. We found that water kefir communities were generally stable at the species-level, with only rare bacterial species replaced over long timescales (more than 2 years). While we observed little strain-level diversity with few strain replacements over long timescales, closely related strains exhibited variation in accessory gene content, often encoded on plasmids, particularly those involved in ecologically meaningful functions such as sugar utilization pathways and phage defense systems. We hypothesise that these genomic variations could reflect the adaptations of strains to different sugars and phages. Consistent with this, we observed a diverse array of phages, many likely originating from the unique household environments. By documenting the genomic landscape of microbial species, strains, plasmids, and phages, this study advances our understanding of the diversity and dynamics of microbial communities in fermented foods. Furthermore, our course material is publicly available and offers a blueprint for bridging the gap between teaching and research, inspiring the next generation of scientists to unravel the complexities of microbial ecosystems. Graphical abstract
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".