Activity-Independent Discovery of Secondary Metabolites Using Chemical Elicitation and Cheminformatic Inference
Bibliographic record
Abstract
Most existing antibiotics were discovered through screens of environmental microbes, particularly the streptomycetes, for the capacity to prevent the growth of pathogenic bacteria. This "activity-guided screening" method has been largely abandoned because it repeatedly rediscovers those compounds that are highly expressed during laboratory culture. Most of these metabolites have already been biochemically characterized. However, the sequencing of streptomycete genomes has revealed a large number of "cryptic" secondary metabolic genes that are either poorly expressed in the laboratory or that have biological activities that cannot be discovered through standard activity-guided screens. Methods that reveal these uncharacterized compounds, particularly methods that are not biased in favor of the highly expressed metabolites, would provide direct access to a large number of potentially useful biologically active small molecules. To address this need, we have devised a discovery method in which a chemical elicitor called Cl-ARC is used to elevate the expression of cryptic biosynthetic genes. We show that the resulting change in product yield permits the direct discovery of secondary metabolites without requiring knowledge of their biological activity. We used this approach to identify three rare secondary metabolites and find that two of them target eukaryotic cells and not bacterial cells. In parallel, we report the first paired use of cheminformatic inference and chemical genetic epistasis in yeast to identify the target. In this way, we demonstrate that oxohygrolidin, one of the eukaryote-active compounds we identified through activity-independent screening, targets the V1 ATPase in yeast and human cells and secondarily HSP90.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".