Phylogenetic Reconstruction of the Legionella pneumophila Philadelphia-1 Laboratory Strains through Comparative Genomics
Bibliographic record
Abstract
Over 20 years ago, two groups independently domesticated Legionella pneumophila from a clinical isolate of bacteria collected during the first recognized outbreak of Legionnaires' disease (at the 1976 American Legion's convention in Philadelphia). These two laboratory strains, JR32 and Lp01, along with their derivatives, have been disseminated to a number of laboratories around the world and form the cornerstone of much of the research conducted on this important pathogen to date. Nevertheless, no exhaustive examination of the genetic distance between these strains and their clinical progenitor has been performed thus far. Such information is of paramount importance for making sense of several phenotypic differences observed between these strains. As environmental replication of L. pneumophila is thought to exclusively occur within natural protozoan hosts, retrospective analysis of the domestication and axenic culture of the Philadelphia-1 progenitor strain by two independent groups also provides an excellent opportunity to uncover evidence of adaptation to the laboratory environment. To reconstruct the phylogenetic relationships between the common laboratory strains of L. pneumophila Philadelphia-1 and their clinical ancestor, we performed whole-genome Illumina resequencing of the two founders of each laboratory lineage: JR32 and Lp01. As expected from earlier, targeted studies, Lp01 and JR32 contain large deletions in the lvh and tra regions, respectively. By sequencing additional strains derived from Lp01 (Lp02 and Lp03), we retraced the phylogeny of these strains relative to their reported ancestor, thereby reconstructing the evolutionary dynamics of each laboratory lineage from genomic data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".