Defining criteria to interpret multilocus variable-number tandem repeat analysis to aid Clostridium difficile outbreak investigation
Bibliographic record
Abstract
PFGE is currently the North American standard for surveillance for Clostridium difficile but lacks discriminatory power to aid outbreak investigation. A further limitation to PFGE is the high baseline rate of the epidemic North American pulsotype (NAP) 1 strain in hospitals. Multilocus variable-number tandem repeat analysis (MLVA) appears to have superior discriminatory power but criteria to define clonality have not been set. We conducted surveillance for toxin-positive C. difficile infection (CDI) at a single academic health sciences centre between September 2009 and April 2010. Seventy-four patient specimens resulting in 86 discrete CDI episodes were subjected to PFGE and MLVA. Results were analysed using Bionumerics software to generate phylogenetic trees and coupled to patient demographic data. Amongst the NAP1 strains, two distinct clusters were identified by MLVA using 90 % similarity as a cut-off by Manhattan distance-based clustering, four clusters using 95 % and seven clusters using 97 %. Population analysis conducted on multiple colonies (n = 25) demonstrated that 1-3 % difference in MLVA types was typical for a single individual. Typing was also conducted in the context of institutional outbreaks (n = 42, three outbreaks) in order to determine clusters within the NAP1 strain. By combining longitudinal surveillance with epidemiological information, single specimen population analysis and typing in the context of institutional outbreaks, we conclude that the use of the Manhattan distance-based clustering with a cut-off of 95-97 % is capable of distinguishing outbreak clones from sporadic isolates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".