Rapid Determination of<i>Escherichia coli</i>O157:H7 Lineage Types and Molecular Subtypes by Using Comparative Genomic Fingerprinting
Bibliographic record
Abstract
In this study, variably absent or present (VAP) regions discovered through comparative genomics experiments were targeted for the development of a rapid, PCR-based method to subtype and fingerprint Escherichia coli O157:H7. Forty-four VAP loci were analyzed for discriminatory power among 79 E. coli O157:H7 strains of 13 phage types (PT). Twenty-three loci were found to maximize resolution among strains, generating 54 separate fingerprints, each of which contained strains of unique PT. Strains from the three previously identified major E. coli O157:H7 lineages, LSPA6-LI, LSPA6-LI/II, and LSPA6-LII, formed distinct branches on a dendrogram obtained by hierarchical clustering of comparative genomic fingerprinting (CGF) data. By contrast, pulsed-field gel electrophoresis (PFGE) typing generated 52 XbaI digestion profiles that were not unique to PT and did not cluster according to O157:H7 lineage. Our analysis identified a subpopulation comprised of 25 strains from a closed herd of cattle, all of which were of PT87 and formed a cluster distinct from all other E. coli O157:H7 strains examined. CGF found five related but unique fingerprints among the highly clonal herd strains, with two dominant subtypes characterized by a shift from the presence of locus fprn33 to its absence. CGF had equal resolution to PFGE typing but with greater specificity, generating fingerprints that were unique among phenotypically related E. coli O157:H7 lineages and PT. As a comparative genomics typing method that is amenable for use in high-throughput platforms, CGF may be a valuable tool in outbreak investigations and strain characterization.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".