Improved Identification of Epidemiologically Related Strains of <i>Salmonella enterica</i> by use of a Fusion Algorithm Based on Pulsed-Field Gel Electrophoresis and Multiple-Locus Variable-Number Tandem-Repeat Analysis
Bibliographic record
Abstract
Pulsed-field gel electrophoresis (PFGE) and multiple-locus variable-number tandem-repeat analysis (MLVA) are used to assess genetic similarity between bacterial strains. There are cases, however, when neither of these methods quantifies genetic variation at a level of resolution that is well suited for studying the molecular epidemiology of bacterial pathogens. To improve estimates based on these methods, we propose a fusion algorithm that combines the information obtained from both PFGE and MLVA assays to assess epidemiological relationships. This involves generating distance matrices for PFGE data (Dice coefficients) and MLVA data (single-step stepwise-mutation model) and modifying the relative distances using the two different data types. We applied the algorithm to a set of Salmonella enterica serovar Typhimurium isolates collected from a wide range of sampling dates, locations, and host species. All three classification methods (PFGE only, MLVA only, and fusion) produced a similar pattern of clustering relative to groupings of common phage types, with the fusion results being slightly better. We then examined a group of serovar Newport isolates collected over a limited geographic and temporal scale and showed that the fusion of PFGE and MLVA data produced the best discrimination of isolates relative to a collection site (farm). Our analysis shows that the fusion of PFGE and MLVA data provides an improved ability to discriminate epidemiologically related isolates but provides only minor improvement in the discrimination of less related isolates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".