Genomic variation across a clinical <i>Cryptococcus</i> population linked to disease outcome
Bibliographic record
Abstract
Abstract Cryptococcus neoformans is the causative agent of cryptococcosis, a disease with poor patient outcomes, accounting for approximately 180,000 deaths each year. Patient outcomes may be impacted by the underlying genetics of the infecting isolate; however, our current understanding of how genetic diversity contributes to clinical outcomes is limited. Here, we leverage clinical, in vitro growth and genomic data for 284 C. neoformans isolates to identify clinically relevant pathogen variants within a population of clinical isolates from patients with HIV-associated cryptococcosis in Malawi. Through a genome-wide association study (GWAS) approach, we identify variants associated with fungal burden and growth rate. We also find both small and large-scale variation, including aneuploidy, associated with alternate growth phenotypes, which may impact the course of infection. Genes impacted by these variants are involved in transcriptional regulation, signal transduction, glycosylation, sugar transport, and glycolysis. We show that growth within the CNS is reliant upon glycolysis in an animal model, and likely impacts patient mortality, as CNS yeast burden likely modulates patient outcome. Additionally, we find genes with roles in sugar transport are enriched in regions under selection in specific lineages of this clinical population. Further, we demonstrate that genomic variants in two genes identified by GWAS impact virulence in animal models. Our approach identifies links between genetic variation in C. neoformans and clinically relevant phenotypes and animal model pathogenesis; shedding light on specific survival mechanisms within the CNS and identifying pathways involved in yeast persistence. Importance Infection outcomes for cryptococcosis, most commonly caused by C. neoformans , are influenced by host immune responses, as well as host and pathogen genetics. Infecting yeast isolates are genetically diverse; however, we lack a deep understanding of how this diversity impacts patient outcomes. To better understand both clinical isolate diversity and how diversity contributes to infection outcome, we utilize a large collection of clinical C. neoformans samples, isolated from patients enrolled in a clinical trial across 3 hospitals in Malawi. By combining whole-genome sequence data, clinical data, and in vitro growth data, we utilize genome-wide association approaches to examine the genetic basis of virulence. Genes with significant associations display virulence attributes in both murine and rabbit models, demonstrating that our approach can identify potential links between genetic variants and patho-biologically significant phenotypes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".