Rapid antimicrobial resistance prediction and identification of Mycobacterium tuberculosis complex by the use of whole genome sequencing on patient sputa
Bibliographic record
Abstract
Tuberculosis (TB) is primarily a respiratory disease caused by the bacterium Mycobacterium tuberculosis (MTB) and accounted for the deaths of 1.6 million people in 2021. TB is typically a treatable disease but drug resistance has become a major public health threat since the 1990s. Current drug susceptibility testing of MTB relies on culture which is slow, labour intensive, and requires specialized infrastructure that may not be available in some regions. Rapid detection of MTB from patient sputum using whole genome sequencing could provide high discriminatory information, while reducing diagnostic turn-around-times thereby preventing costly and ineffective treatments. The aim of this study was to develop a culture-free genomic method to identify MTB and predict drug resistance from sputum using whole genome sequencing. A validation study was done in two stages: (1) MTB-negative sputum spiked with Mycobacterium bovis BCG, (2) MTB positive sputum. Sputum is the primary clinical specimen for MTB testing but presents challenges due to the presence of both high host and microbial DNA compared to MTB. Sputum was decontaminated using standard methods to liquefy sample and reduce host and bacterial presence. Samples were enzymatically treated to degrade host DNA. Prior to sequencing, extracts were subjected to two amplification methods: an in-house developed multiplex-PCR targeting known MTB resistance markers and a random approach using GC-rich primers. Sequences obtained from Illumina MiSeq and Oxford Nanopore Technologies (ONT) were subjected to quality analysis in Galaxy (version v20.01) before being submitted to bioinformatics pipelines: Kraken2, Mykrobe Predictor and BioHansel which were used to assess bacterial and human presence, predict antimicrobial resistance (AMR) determinants and species identification, respectively. Results indicated that amplification methods improved both DNA concentrations and genome coverage by increasing mycobacterial DNA abundance. The optimized protocol performed best with ONT, generating higher mycobacterial genomic coverage and depth which improved AMR predictions. The finalized protocol provides promising steps forward to deploying rapid diagnostics for MTB directly from sputum.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".