Discovery of the missing cytochrome P450 monooxygenase cyclases that conclude glyceollin biosynthesis in soybean
Bibliographic record
Abstract
Abstract Glyceollins are isoflavonoid-derived metabolites produced by soybean that hold great promise in improving human and animal health due to their antimicrobial, and other medicinal properties. They play important roles in agriculture by defending soybean against one of its most destructive pathogens, Phytophthora sojae . Longstanding research efforts have focused on improving accessibility to glyceollins, yet chemical synthesis remains uneconomical. The fact that some of the key genes involved in the final step of glyceollin biosynthesis have not been identified, engineering the accumulation of these important compounds in microbes is not yet possible. Although the activity of a P450 cyclase was inferred to catalyze the final committed step in glyceollin biosynthesis forty years ago, the enzyme in question has never been conclusively identified. This study reports, for the first time, the identification of three cytochrome P450 monooxygenase cyclases that catalyze the final steps of glyceollin biosynthesis. Utilizing P. sojae -soybean transcriptome data, along with genome mining tools and co-expression network analysis, we have identified 16 candidate glyceollin synthases (GmGS). Heterologous expression of these candidate genes in yeast, coupled with in vitro enzyme assays, enabled us to discover three enzymes capable of producing two glyceollin isomers. GmGS11A and GmGS11B catalyzed the conversion of glyceollidin to glyceollin I, whereas GmGS13A converted glyceocarpin to glyceollin III. The functionality of these candidates was further confirmed in planta through gene silencing and overexpression in soybean hairy roots. This groundbreaking study not only contributes to the understanding of glyceollin biosynthesis, but also demonstrates a new synthetic biology strategy that could potentially be scaled up to produce valuable molecules for crop and disease management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".