Evolution of a secondary metabolic pathway from primary metabolism: shikimate and quinate biosynthesis in plants
Bibliographic record
Abstract
The shikimate pathway synthesizes aromatic amino acids essential for protein biosynthesis. Shikimate dehydrogenase (SDH) is a central enzyme of this primary metabolic pathway, producing shikimate. The structurally similar quinate is a secondary metabolite synthesized by quinate dehydrogenase (QDH). SDH and QDH belong to the same gene family, which diverged into two phylogenetic clades after a defining gene duplication just prior to the angiosperm/gymnosperm split. Non-seed plants that diverged before this duplication harbour only a single gene of this family. Extant representatives from the chlorophytes (Chlamydomonas reinhardtii), bryophytes (Physcomitrella patens) and lycophytes (Selaginella moellendorfii) encoded almost exclusively SDH activity in vitro. A reconstructed ancestral sequence representing the node just prior to the gene duplication also encoded SDH activity. Quinate dehydrogenase activity was gained only in seed plants following gene duplication. Quinate dehydrogenases of gymnosperms, represented here by Pinus taeda, may be reminiscent of an evolutionary intermediate since they encode equal SDH and QDH activities. The second copy in P. taeda maintained specificity for shikimate similar to the activity found in the angiosperm SDH sister clade. The codon for a tyrosine residue within the active site displayed a signature of positive selection at the node defining the QDH clade, where it changed to a glycine. Replacing the tyrosine with a glycine in a highly shikimate-specific angiosperm SDH was sufficient to gain some QDH function. Thus, very few mutations were necessary to facilitate the evolution of QDH genes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".