Typhi Mykrobe: fast and accurate lineage identification and antimicrobial resistance genotyping directly from sequence reads for the typhoid fever agent Salmonella Typhi
Bibliographic record
Abstract
BACKGROUND: Typhoid fever results from systemic infection with Salmonella enterica serovar Typhi (Typhi) and causes 10 million illnesses annually. Disease control relies on prevention (water, sanitation, and hygiene interventions or vaccination) and effective antimicrobial treatment. Antimicrobial-resistant (AMR) Typhi lineages have emerged and become established in many parts of the world. Knowledge of local pathogen populations informed by genomic surveillance, including of lineages (defined by the GenoTyphi scheme) and AMR determinants, is increasingly used to inform local treatment guidelines and to inform vaccination strategy. Current tools for genotyping Typhi require multiple read alignment or assembly steps and have not been validated for analysis of data generated with Oxford Nanopore Technologies (ONT) long-read sequencing devices. Here, we introduce Typhi Mykrobe, a command line software tool for rapid genotyping of Typhi lineages, AMR determinants, and plasmid replicons direct from sequencing reads. RESULTS: We validated Typhi Mykrobe lineage genotyping by comparison with the current standard read mapping-based approach and demonstrated 99.8% concordance across nearly 13,000 genomes sequenced with Illumina platforms. For the few isolates with discordant calls, we show that Typhi Mykrobe results are better supported by the evidence from raw sequence read data than the results generated using the mapping-based approach. We also demonstrate 99.9% concordance for detection of AMR determinants compared with the current standard assembly-based approach, with similar results for plasmid marker detection. Typhi Mykrobe predicts clinical resistance categorization (S/I/R) for eight drug classes, and we show strong agreement with phenotypic categorizations generated from reference laboratory minimum inhibitory concentration (MIC) data for n = 1572 Illumina-sequenced isolates (> 99% agreement within one doubling dilution). We show strong concordance (> 96% for genotype and > 98% for AMR and plasmid) between calls made from ONT reads and those made from Illumina reads for isolates sequenced on both platforms (n = 93 genomes). Typhi Mykrobe takes less than a minute per sample and is available at https://github.com/typhoidgenomics/genotyphi . CONCLUSIONS: Typhi Mykrobe provides rapid and sensitive genotyping of Typhi genomes direct from Illumina and ONT reads, although lower accuracy was observed for R9 ONT data. It demonstrated accurate assignment of GenoTyphi lineage, detection of AMR determinants and prediction of corresponding AMR phenotypes, and identification of plasmid replicons.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".