Improved Short-Sequence-Repeat Genotyping of Mycobacterium avium subsp. paratuberculosis by Using Matrix-Assisted Laser Desorption Ionization–Time of Flight Mass Spectrometry
Bibliographic record
Abstract
Accurate sequence analysis of mononucleotide repeat regions is difficult, complicating the use of short sequence repeats (SSRs) as a tool for bacterial strain discrimination. Although multiple SSR loci in the genome of Mycobacterium avium subsp. paratuberculosis allow genotyping of M. avium subsp. paratuberculosis isolates with high discriminatory power, further characterization of the most discriminatory loci is limited due to inherent difficulties in sequencing mononucleotide repeats. Here, a method was evaluated using matrix-assisted laser desorption ionization-time of flight mass spectrometry (MALDI-TOF MS) as an alternative to Sanger sequencing to further differentiate the dominant mycobacterial interspersed repetitive-unit (MIRU)-variable-number tandem-repeat (VNTR) M. avium subsp. paratuberculosis type (n = 37) in Canadian dairy herds by targeting a highly discriminatory mononucleotide SSR locus. First, PCR-amplified DNA was digested with two restriction enzymes to yield a sufficiently small fragment containing the SSR locus. Second, MALDI-TOF MS was performed to identify the mass, and thus repeat length, of the target. Sufficiently intense, discriminating spectra were obtained to determine repeat lengths up to 15, an improvement over the limit of 11 using traditional sequencing techniques. Comparison to synthetic oligonucleotides and Sanger sequencing results confirmed a valid and reproducible assay that increased discrimination of the dominant M. avium subsp. paratuberculosis MIRU-VNTR type. Thus, MALDI-TOF MS was a reliable, fast, and automatable technique to accurately resolve M. avium subsp. paratuberculosis genotypes based on SSRs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".