Comparative genomics reveals intra and inter species variation in the pathogenic fungus <i>Batrachochytrium dendrobatidis</i>
Bibliographic record
Abstract
Abstract The Global Panzootic Lineage (GPL) of the amphibian pathogen Batrachochytrium dendrobatidis ( Bd ) has been described as a main driver of amphibian extinctions on nearly every continent. Near complete genomes of three Bd -GPL strains have enabled studies of the pathogen but the genomic features that set Bd- GPL apart from other B. dendrobatidis lineages is not well understood due to a lack of high-quality genome assemblies and annotations from other lineages. We used Oxford Nanopore Technologies (ONT) DNA sequencing to assemble high-quality genomes of three Bd -BRAZIL isolates and one non-pathogen outgroup species Polyrhizophydium stewartii ( Ps ) strain JEL0888 and compared these to genomes of previously sequenced Bd- GPL strains. The Bd -BRAZIL assemblies range in size between 22.0 and 26.1 Mb and encode 8495-8620 protein-coding genes for each strain. A pangenome is defined as all the genes in a species including the core genome, genes found in every strain, and the accessory genome, genes found in only some strains (Brockhurst et al. 2019). To date, a comprehensive analysis identifying the core and accessory genes within B. dendrobatidis has not been conducted. Furthermore, while previous studies have examined the gene transcription profiles of Bd -GPL and Bd -BRAZIL strains (McDonald et al. 2020), they do not account for the genomic differences between these strains. Our pangenome analysis provides insight into shared and lineage-specific gene content and how B. dendrobatidis genotype affects recovery of RNAseq transcripts from different strains. We hypothesize that gene content differences exist between the B. dendrobatidis lineages and genomic differences, such as gene family expansions or gene sequence variation, affect alignment and enumeration of transcriptomic data when relying on a single reference genome. The pangenome analysis revealed a core genome consisting of 6278 conserved gene families, and an accessory genome with 202 Bd -BRAZIL and 172 Bd -GPL specific gene families. We discovered gene copy number differences in five pathogenicity gene families: M36 Peptidase, Crinkler Necrosis Genes (CRN), Aspartyl Peptidase, Carbohydrate-Binding Module-18 (CBM18), and S41 Protease, between Bd- BRAZIL and Bd- GPL strains. However, none of the five families were expanded in Bd -GPL compared to Bd -BRAZIL strains. Comparison between the Batrachochytrium genus and two closely related non-pathogenic saprophytic chytrids identified differences in sequence and protein domain counts. We further test these new Bd -BRAZIL genomes to assess their utility as reference genomes for transcriptome alignment and analysis. Our analysis examines the genomic variation between strains in Bd -BRAZIL and Bd -GPL and offers insights into the application of these genomes as reference genomes for future studies. Significance The geographically defined enzootic populations of amphibian pathogen Batrachochytrium dendrobatidis harbor gene context variation revealed in pan-genome analyses. Long read sequencing is required to fully capture this diversity as some recently duplicated and potential virulence gene families are undercounted in short-read only genome assemblies. This genetic variation can impact estimates of gene expression differences between strains if a single genome reference is used. It is necessary to consider the pan-genome diversity of the multiple lineages of this important amphibian pathogen and perhaps other fungal pathogens when engaging in studies of adaptation, virulence, and comparative biology of a species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".