P-397. Rapid Prediction of Genetic Relatedness of <i>Escherichia coli</i> Direct from Critical Care Samples Using Metagenomic Sequencing Paired with Neighbour Typing
Bibliographic record
Abstract
Abstract Background Determining relatedness of bacterial pathogens within hospitals is important to rule out transmission events, but current approaches are typically slow and/or resource intensive. We sought to assess the utility of rapid long-read sequencing with nearest neighbour-typing for identifying relatedness direct from specimens. Figure 1. Plot of true and predicted genetic distances for E. coli. (a) All comparisons. (B) Only concordant comparisons. (C) All comparisons with LS&lt;0.5. (D) All comparisons with LS&gt;0.5. Green data points are predictions based on concordant calls, and orange data points are predictions based on discordant calls. Methods We sequenced remnant urine (n=71) and respiratory (n=7) samples with confirmed Escherichia coli (n=78) using ONT MinION, from critical care units at two large tertiary care institutions in Ontario, Canada. The resistance RASE method was used to predict the nearest neighbour for all samples. Using the nearest neighbours and the known reference database phylogeny, we predicted genetic (pairwise) distances (percentage). We then compared the predicted genetic distances with the true genetic distances. Plots visualizing the true and predicted genetic distances were generated and linear models performed. Genetic distance predictions were categorized as: (1) concordant (matching predicted and true sequence types) vs. discordant; and (2) high confidence in prediction (lineage score [LS] &gt;0.5) vs. low confidence in prediction (LS &lt; 0.5). Results For all specimens, predicted genetic distances demonstrated a modest but significant correlation (R2=0.52, p&lt; 0.001) with true genetic distances (Fig 1A). Predictions of genetic distance were most accurate when the predicted sequence type matched the true sequence type (R2=0.99, p&lt; 0.001) (Fig 1B). Using the lineage score improved the accuracy of genetic distance predictions (R2=0.744, p&lt; 0.001) (Fig 1C-D). Using a genetic distance cut-off of 0.002 to denote unrelated (non-transmitted) events, as determined by the frequency distribution of genetic distances, a lineage score informed approach could identify genetically unrelated E. coli with a specificity of 0.99, sensitivity of 0.96, and positive predictive value of 0.99. Conclusion Rapid predictions of genetic-relatedness can be made directly from specimens using long-read sequencing paired with neighbour-typing. We demonstrate that in high confidence calls, the potential for related transmission between infections can be effectively ruled out. Disclosures Allison McGeer, MD, AstraZeneca: Honoraria|GSK: Honoraria|Merck: Honoraria|Moderna: Honoraria|Novavax: Honoraria|Pfizer: Grant/Research Support|Pfizer: Honoraria|Roche: Honoraria|Seqirus: Grant/Research Support|Seqirus: Honoraria
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".