Activity-Independent Discovery of Secondary Metabolites Using Chemical Elicitation and Cheminformatic Inference
Bibliographic record
Abstract
Most existing antibiotics were discovered through screens of environmental microbes, particularly the streptomycetes, for the capacity to prevent the growth of pathogenic bacteria. This "activity-guided screening" method has been largely abandoned because it repeatedly rediscovers those compounds that are highly expressed during laboratory culture. Most of these metabolites have already been biochemically characterized. However, the sequencing of streptomycete genomes has revealed a large number of "cryptic" secondary metabolic genes that are either poorly expressed in the laboratory or that have biological activities that cannot be discovered through standard activity-guided screens. Methods that reveal these uncharacterized compounds, particularly methods that are not biased in favor of the highly expressed metabolites, would provide direct access to a large number of potentially useful biologically active small molecules. To address this need, we have devised a discovery method in which a chemical elicitor called Cl-ARC is used to elevate the expression of cryptic biosynthetic genes. We show that the resulting change in product yield permits the direct discovery of secondary metabolites without requiring knowledge of their biological activity. We used this approach to identify three rare secondary metabolites and find that two of them target eukaryotic cells and not bacterial cells. In parallel, we report the first paired use of cheminformatic inference and chemical genetic epistasis in yeast to identify the target. In this way, we demonstrate that oxohygrolidin, one of the eukaryote-active compounds we identified through activity-independent screening, targets the V1 ATPase in yeast and human cells and secondarily HSP90.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".