Vibrio cholerae lineage and pangenome diversity vary geographically across Bangladesh over 1 year
Bibliographic record
Abstract
Cholera is an acute diarrhoeal disease caused by Vibrio cholerae . It remains a major public health challenge worldwide, and particularly in the endemic region around the Bay of Bengal. Over decadal time scales, one lineage typically dominates and spreads in global pandemic waves. However, it remains unclear to what extent diverse lineages co-circulate during a single outbreak. Defining the pool of diversity over finer time-scales is important because the selective pressures that impact V. cholerae , namely antibiotics and phages, are dynamic on these scales. To study the nationwide diversity of V. cholerae , we long-read sequenced 273 V . cholerae genomes from seven hospitals over 1 year (2018) in Bangladesh. Four major V. cholerae lineages were identified: three known lineages, BD-1, BD-2a and BD-2b, and a novel lineage that we call BD-3. In 2022, BD-1 caused a large cholera outbreak in Dhaka, at which point it had replaced BD-2 as the most common lineage in Bangladesh. We show that, in 2018, BD-1 was already predominant in the five northern regions, including Dhaka, consistent with an origin from northern India. By contrast, we observed a higher diversity of lineages in the two southern regions near the coast. The four lineages differed in pangenome content, including integrative and conjugative elements (ICEs) and genes involved in resistance to bacteriophages and antibiotics. Notably, BD-2a lacked an ICE and is predicted to be more sensitive to phages and antibiotics, yet persisted throughout the sampling period. Genes previously associated with antibiotic resistance in V. cholerae isolated from Bangladesh in the prior decade were entirely absent from all lineages in 2018–2019, suggesting shifting costs and benefits of encoding these genes. Our results highlight the diverse nature of the V. cholerae pangenome and geographic structure within a single outbreak season. This diversity provides the raw material for adaptation to antibiotics, phages and other selective pressures.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".