MétaCan
Menu
Back to cohort
Record W4365135538 · doi:10.1099/mgen.0.000979

Comparing genomic variant identification protocols for Candida auris

2023· article· en· W4365135538 on OpenAlexfundno aff
Li Xiao, José F. Muñoz, Lalitha Gade, Silvia Argimón, Marie‐Elisabeth Bougnoux, Jolene R. Bowers, Nancy A. Chow, Isabel Cuesta, Rhys A. Farrer, Corinne Maufrais, Juan Monroy-Nieto, Dibyabhaba Pradhan, Jessie Uehling, Duong Vu, Corin Yeats, David M. Aanensen, Christophe d’Enfert, David M. Engelthaler, David W. Eyre, Matthew C. Fisher, Ferry Hagen, Wieland Meyer, Gagandeep Singh, Ana Alastruey‐Izquierdo, Anastasia P. Litvintseva, Christina A. Cuomo

Bibliographic record

VenueMicrobial Genomics · 2023
Typearticle
Languageen
FieldMedicine
TopicAntifungal resistance and susceptibility
Canadian institutionsnot available
FundersCommissariat Général à l'InvestissementCenters for Disease Control and PreventionAgence Nationale de la RechercheNational Institute of Allergy and Infectious DiseasesDivision of Intramural Research, National Institute of Allergy and Infectious DiseasesBroad InstituteWellcome TrustMedical Research CouncilCanadian Institute for Advanced ResearchNational Institutes of HealthU.S. Department of Health and Human Services
KeywordsBiologySingle-nucleotide polymorphismPhylogenetic treeCladeComputational biologyGeneticsConcordanceGenome-wide association studySNPPopulationGeneMedicineGenotype

Abstract

fetched live from OpenAlex

Genomic analyses are widely applied to epidemiological, population genetic and experimental studies of pathogenic fungi. A wide range of methods are employed to carry out these analyses, typically without including controls that gauge the accuracy of variant prediction. The importance of tracking outbreaks at a global scale has raised the urgency of establishing high-accuracy pipelines that generate consistent results between research groups. To evaluate currently employed methods for whole-genome variant detection and elaborate best practices for fungal pathogens, we compared how 14 independent variant calling pipelines performed across 35 Candida auris isolates from 4 distinct clades and evaluated the performance of variant calling, single-nucleotide polymorphism (SNP) counts and phylogenetic inference results. Although these pipelines used different variant callers and filtering criteria, we found high overall agreement of SNPs from each pipeline. This concordance correlated with site quality, as SNPs discovered by a few pipelines tended to show lower mapping quality scores and depth of coverage than those recovered by all pipelines. We observed that the major differences between pipelines were due to variation in read trimming strategies, SNP calling methods and parameters, and downstream filtration criteria. We calculated specificity and sensitivity for each pipeline by aligning three isolates with chromosomal level assemblies and found that the GATK-based pipelines were well balanced between these metrics. Selection of trimming methods had a greater impact on SAMtools-based pipelines than those using GATK. Phylogenetic trees inferred by each pipeline showed high consistency at the clade level, but there was more variability between isolates from a single outbreak, with pipelines that used more stringent cutoffs having lower resolution. This project generated two truth datasets useful for routine benchmarking of C. auris variant calling, a consensus VCF of genotypes discovered by 10 or more pipelines across these 35 diverse isolates and variants for 2 samples identified from whole-genome alignments. This study provides a foundation for evaluating SNP calling pipelines and developing best practices for future fungal genomic studies.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.012
metaresearch head score (Gemma)0.031
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.012
Threshold uncertainty score0.063

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0120.031
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0020.002
Science and technology studies0.0010.001
Scholarly communication0.0020.001
Open science0.0010.002
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.057
GPT teacher head0.325
Teacher spread0.268 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations15
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueMicrobial GenomicsSame topicAntifungal resistance and susceptibilityFrench-language works237,207