Comparative Genomics Reveal That Host-Innate Immune Responses Influence the Clinical Prevalence of Legionella pneumophila Serogroups
Bibliographic record
Abstract
Legionella pneumophila is the primary etiologic agent of legionellosis, a potentially fatal respiratory illness. Amongst the sixteen described L. pneumophila serogroups, a majority of the clinical infections diagnosed using standard methods are serogroup 1 (Sg1). This high clinical prevalence of Sg1 is hypothesized to be linked to environmental specific advantages and/or to increased virulence of strains belonging to Sg1. The genetic determinants for this prevalence remain unknown primarily due to the limited genomic information available for non-Sg1 clinical strains. Through a systematic attempt to culture Legionella from patient respiratory samples, we have previously reported that 34% of all culture confirmed legionellosis cases in Ontario (n = 351) are caused by non-Sg1 Legionella. Phylogenetic analysis combining multiple-locus variable number tandem repeat analysis and sequence based typing profiles of all non-Sg1 identified that L. pneumophila clinical strains (n = 73) belonging to the two most prevalent molecular types were Sg6. We conducted whole genome sequencing of two strains representative of these sequence types and one distant neighbour. Comparative genomics of the three L. pneumophila Sg6 genomes reported here with published L. pneumophila serogroup 1 genomes identified genetic differences in the O-antigen biosynthetic cluster. Comparative optical mapping analysis between Sg6 and Sg1 further corroborated this finding. We confirmed an altered O-antigen profile of Sg6, and tested its possible effects on growth and replication in in vitro biological models and experimental murine infections. Our data indicates that while clinical Sg1 might not be better suited than Sg6 in colonizing environmental niches, increased bloodstream dissemination through resistance to the alternative pathway of complement mediated killing in the human host may explain its higher prevalence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".