Ciprofloxacin resistance in <i>Klebsiella pneumoniae</i> : phenotype prediction from genotype and global distribution of resistance determinants
Bibliographic record
Abstract
ABSTRACT BACKGROUND Ciprofloxacin resistant Klebsiella pneumoniae is common or emerging in many geographies, and knowledge of local resistance rates is important for empirical therapy. Whilst there are known K. pneumoniae ciprofloxacin resistance determinants, there is a lack of systematic data on the effect of determinants, alone and in combination, and there are no publicly accessible tools for predicting resistance from whole genome sequence data. METHODS The KlebNET-GSP AMR Genotype-Phenotype Group aggregated a matched genotype-phenotype dataset of n=12,167 K. pneumoniae species complex ( Kp SC) isolates from 27 countries between 2001-2021. We developed a rules-based classifier to predict ciprofloxacin resistance by categorizing the number of quinolone resistance determining regions mutations in gyrA and parC , the number of plasmid-mediated quinolone resistance genes, and the presence/absence of aac(6ʹ)-Ib-cr (which can acetylate ciprofloxacin). Predictive performance was assessed using the discovery dataset, for which we re-phenotyped discrepant isolates; and validated using externally contributed datasets (n=7,030 Kp SC isolates). RESULTS The rules-based classifier predicted R vs S/I with categorical agreement, sensitivity, and specificity >96%, and major/very major error rates <4%. Performance was similar across diverse Kp SC sources (human, animal, other), species, and intra-species lineages. External validation of the classifier yielded overall 93.12% categorical agreement [95% confidence interval (CI), 92.50-93.74%], 8.65% major errors [95% CI, 7.34-9.97%], and 6.20% very major errors [95% CI, 5.51-6.90%]. We implemented the classifier in Kleborate, a command-line tool that is integrated into the Pathogenwatch web platform. Using this to assess the global distribution of ciprofloxacin resistance determinants in Kp SC genomes available in Pathogenwatch (n=31,319, from 109 countries between years 2000-2023), we observed a significant positive association between national quinolone consumption rates and predicted ciprofloxacin resistance (R 2 =0.20, p=0.004). CONCLUSIONS Ciprofloxacin resistance phenotypes can be reasonably predicted from genotypes, which is sufficient for informing surveillance. However, unexplained resistance remains and accuracy is insufficient for clinical applications. We demonstrate the value of aggregating genotype-phenotype data to explore resistance mechanisms and develop predictors, but highlight complexities in combining phenotype data from different assays and standards.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".