Paired-omics-based exploration and characterization of biosynthetic diversity in lichenized fungi
Bibliographic record
Abstract
The increasing demand for novel drug leads requires bioprospecting non-model taxa. Comparative genomics and correlative omics are a fast and efficient method for linking bioactive but genetically orphan natural products to their biosynthetic gene clusters (BGCs) and identifying potentially novel drug leads. Here we implement these approaches for the first systematic comparison of the BGC diversity in lichen-forming fungi (LFF) (comprising 20% of known fungi), prolific but underutilized producers of bioactive natural products. We first identified BGCs from all publicly available LFF genomes (111), encompassing 71 fungal genera and 23 families, and generated BGC similarity networks of each class. We recovered 5,541 BGCs grouped into 4,464 gene cluster families. We used mass spectrometry (MS) and correlative metabolomics to link five MS-identified metabolites - alectoronic acid, alpha-collatolic acid, evernic acid, stenosporic acid and perlatolic acid - to their putative BGCs. We subsequently used MS on an additional 80 species to explore the taxonomic breadth of common lichen compounds, uncovering a strong pattern between specific families and secondary metabolites. We found that (1) ~98% of the BGCs in LFF are putatively novel (uncharacterized to date), (2) lichen metabolic profiles contain a plethora of unidentified metabolites and (3) ribosomal peptide-related BGCs constitute about 20% of the LFF BGC landscape. Our study provides comprehensive insights into the BGC landscape of LFFs, highlighting unique, widespread and previously uncharacterized BGCs. We anticipate that the approach we describe will serve as a baseline for leveraging biosynthetic research in non-model organisms, inspiring further investigations into microbial dark matter.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".