Comprehensive cross‐genome survey and phylogeny of glycoside hydrolase family 16 members reveals the evolutionary origin of <scp>EG</scp>16 and <scp>XTH</scp> proteins in plant lineages
Bibliographic record
Abstract
Carbohydrate-active enzymes (CAZymes) are central to the biosynthesis and modification of the plant cell wall. An ancient clade of bifunctional plant endo-glucanases (EG16 members) was recently revealed and proposed to represent a transitional group uniting plant xyloglucan endo-transglycosylase/hydrolase (XTH) gene products and bacterial mixed-linkage endo-glucanases in the phylogeny of glycoside hydrolase family 16 (GH16). To gain broader insights into the distribution and frequency of EG16 and other GH16 members in plants, the PHYTOZOME, PLAZA, NCBI and 1000 PLANTS databases were mined to build a comprehensive census among 1289 species, spanning the broad phylogenetic diversity of multiple algae through recent plant lineages. EG16, newly identified EG16-2 and XTH members appeared first in the green algae. Extant EG16 members represent the early adoption of the β-jellyroll protein scaffold from a bacterial or early-lineage eukaryotic GH16 gene, which is characterized by loop deletion and extension of the N terminus (in EG16-2 members) or C terminus (in XTH members). Maximum-likelihood phylogenetic analysis of EG16 and EG16-2 sequences are directly concordant with contemporary estimates of plant evolution. The lack of expansion of EG16 members into multi-gene families across green plants may point to a core metabolic role under tight control, in contrast to XTH genes that have undergone the extensive duplications typical of cell-wall CAZymes. The present census will underpin future studies to elucidate the physiological role of EG16 members across plant species, and serve as roadmap for delineating the closely related EG16 and XTH gene products in bioinformatic analyses of emerging genomes and transcriptomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".