High-Throughput Short Sequence Typing Schemes for <i>Pseudomonas aeruginosa</i> and <i>Stenotrophomonas maltophilia</i> pure culture and environmental DNA
Bibliographic record
Abstract
Abstract Molecular typing techniques are employed to determine the genetic similarities between bacterial isolates. These methods primarily utilize specific genetic markers or analyze the complete genome sequence of pure bacterial cultures. However, the use of environmental DNA profiling to assess epidemiologic links between patients and their environment has not been explored in depth. This work reports on the development and validation of two High-Throughput Short Sequence Typing (HiSST) schemes targeting the opportunistic pathogens Pseudomonas aeruginosa and Stenotrophomonas maltophilia , along with a modified SM2I medium for specific isolation of S. maltophilia . Our HiSST schemes are based on four discriminative loci for each species and demonstrate high discrimination power, comparable to pairwise whole genomes comparison. Moreover, each scheme includes species-specific PCR primers, enabling precise differentiation from closely related taxa without the need for upstream culture-dependent methods. For example, the primers designed to target the bvgS locus allow to distinguish P. aeruginosa from the very closely related Pseudomonas paraeruginosa sp. nov. The selected loci included in the schemes for P. aeruginosa ( pheT , btuB , sdaA , bvgS ) and for S. maltophilia ( yvoA , glnG , ribA , tycC) , are within the range of 271 to 330 base pairs adapted to massive parallel amplicon sequencing technology. A R-based script implemented in the DADA2 pipeline was assembled to facilitate HiSST analysis for efficient and accurate genotyping of P. aeruginosa and S. maltophilia . The performance of both schemes was demonstrated through in-silico validations, assessments against reference culture collections, and a case study involving environmental samples.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".