MétaCan
Menu
Back to cohort
Record W2999739138 · doi:10.1101/2020.01.10.902056

Rapid and accurate SNP genotyping of clonal bacterial pathogens with BioHansel

2020· preprint· en· W2999739138 on OpenAlexafffund
Geneviève Labbé, Peter Kruczkiewicz, Philip Mabon, James A. Robertson, Justin Schonfeld, Daniel Kein, Marisa A. Rankin, Matthew Gopez, Darian Hole, David Son, Natalie Knox, Chad Laing, Kyrylo Bessonov, Eduardo N. Taboada, Catherine Yoshida, Kim Ziebell, Anil Nichani, Roger P. Johnson, Gary Van Domselaar, John J. Nash

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2020
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Phylogenetic Studies
Canadian institutionsUniversity of ManitobaPublic Health Agency of Canada
FundersGovernment of CanadaPublic Health AgencyPublic Health Agency of Canada
KeywordsGenotypingComputational biologySNP genotypingWhole genome sequencingContigIndelGenomeBiologyGeneticsSingle-nucleotide polymorphismComputer scienceData miningGenotype

Abstract

fetched live from OpenAlex

Abstract BioHansel performs high-resolution genotyping of bacterial isolates by identifying phylogenetically informative single nucleotide polymorphisms (SNPs), also known as canonical SNPs, in whole genome sequencing (WGS) data. The application uses a fast k -mer matching algorithm to map pathogen WGS data to canonical SNPs contained in hierarchically structured schemas and assigns genotypes based on the detected SNP profile. Using modest computing resources, BioHansel efficiently types isolates from raw sequence reads or assembled contigs in a matter of seconds, making it attractive for use by public health, food safety, environmental, and agricultural authorities that wish to apply WGS methodologies for their surveillance, diagnostics, and research programs. BioHansel currently provides canonical SNP genotyping schemas for four prevalent Salmonella serovars—Typhi, Typhimurium, Enteritidis and Heidelberg—as well as a schema for Mycobacterium tuberculosis . Users can also supply their own schemas for genotyping other organisms. BioHansel’s quality assurance system assesses the validity of the genotyping results and can identify low quality data, contaminated datasets, and misidentified organisms. BioHansel is targeted to support surveillance, source attribution, risk assessment, diagnostics, and rapid screening for public health purposes, such as product recalls. BioHansel is an open source application with packages available for PyPI, Conda, and the Galaxy workflow manager. In summary, BioHansel performs efficient, rapid, accurate, and high-resolution classification of bacterial genomes from sequence reads or assembled contigs on standard computing hardware. BioHansel is suitable for use as a general research tool as well as in fully operationalized WGS workflows at the front lines of infectious disease surveillance, diagnostics, and outbreak investigation and response. Impact statement Public health, food safety, environmental, and agricultural authorities are currently engaged in a global effort to incorporate whole genome sequencing technologies into their infectious disease research, surveillance, and outbreak investigation programs. Its widespread adoption, however, has been impeded by two major obstacles: the need for high performance computing to generate results and the expert knowledge required to interpret and communicate those results. BioHansel addresses these limitations by rapidly genotyping pathogens from whole genome sequence data in an accurate, simple, familiar, and easily sharable manner using standard computing resources. BioHansel provides a compact and readily interpretable genotype based on canonical SNP genotyping schemas. BioHansel’s genotyping nomenclature encodes the pathogen’s position in its population structure, which simplifies and facilitates its comparison with actively circulating strains and historical strains. The genotyping information provided by BioHansel can identify points of intervention to prevent the spread of pathogenic bacteria, screen for the presence of priority pathogens, and perform source attribution and risk assessment. Thus, BioHansel serves as a readily accessible and powerful WGS method, implementable on a laptop, for genotyping pathogens to detect, monitor, and control the emergence and spread of infectious disease through surveillance, screening, diagnostics, and outbreak investigation and response activities. Data summary BioHansel is a Python 3 application available as PyPI, Conda Galaxy Tool Shed packages. It is an open source application distributed under the Apache License, Version 2.0. Source code is available at https://github.com/phac-nml/biohansel . The BioHansel user guide is available at https://bio-hansel.readthedocs.io/en/readthedocs/ . Supplementary Materials are available at https://github.com/phac-nml/biohansel-manuscript-supplementary-data . The authors confirm all supporting data, code and protocols have been provided within the article or through supplementary data files.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.009
Threshold uncertainty score0.030

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.003
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0010.001
Scholarly communication0.0020.002
Open science0.0020.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0090.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.211
Teacher spread0.192 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2020
Admission routes2
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicGenomics and Phylogenetic StudiesFrench-language works237,207