GeneSippr: A Rapid Whole-Genome Approach for the Identification and Characterization of Foodborne Pathogens such as Priority Shiga Toxigenic Escherichia coli
Bibliographic record
Abstract
The timely identification and characterization of foodborne bacteria for risk assessment purposes is a key operation in outbreak investigations. Current methods require several days and/or provide low-resolution characterization. Here we describe a whole-genome-sequencing (WGS) approach (GeneSippr) enabling same-day identification of colony isolates recovered from investigative food samples. The identification of colonies of priority Shiga-toxigenic Escherichia coli (STEC) (i.e., serogroups O26, O45, O103, O111, O121, O145 and O157) served as a proof of concept. Genomic DNA was isolated from single colonies and sequencing was conducted on the Illumina MiSeq instrument with raw data sampling from the instrument following 4.5 hrs of sequencing. Modeling experiments indicated that datasets comprised of 21-nt reads representing approximately 4-fold coverage of the genome were sufficient to avoid significant gaps in sequence data. A novel bioinformatic pipeline was used to identify the presence of specific marker genes based on mapping of the short reads to reference sequence libraries, along with the detection of dispersed conserved genomic markers as a quality control metric to assure the validity of the analysis. STEC virulence markers were correctly identified in all isolates tested, and single colonies were identified within 9 hrs. This method has the potential to produce high-resolution characterization of STEC isolates, and whole-genome sequence data generated following the GeneSippr analysis could be used for isolate identification in place of lengthy biochemical characterization and typing methodologies. Significant advantages of this procedure include ease of adaptation to the detection of any gene marker of interest, as well as to the identification of other foodborne pathogens for which genomic markers have been defined.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".