Genome-Wide Identification and Evolutionary and Expression Analyses of MYB-Related Genes in Land Plants
Bibliographic record
Abstract
MYB proteins constitute one of the largest transcription factor families in plants. Recent evidence revealed that MYB-related genes play crucial roles in plants. However, compared with the R2R3-MYB type, little is known about the complex evolutionary history of MYB-related proteins in plants. Here, we present a genome-wide analysis of MYB-related proteins from 16 species of flowering plants, moss, Selaginella, and algae. We identified many MYB-related proteins in angiosperms, but few in algae. Phylogenetic analysis classified MYB-related proteins into five distinct subgroups, a result supported by highly conserved intron patterns, consensus motifs, and protein domain architecture. Phylogenetic and functional analyses revealed that the Circadian Clock Associated 1-like/R-R and Telomeric DNA-binding protein-like subgroups are >1 billion yrs old, whereas the I-box-binding factor-like and CAPRICE-like subgroups appear to be newly derived in angiosperms. We further demonstrated that the MYB-like domain has evolved under strong purifying selection, indicating the conservation of MYB-related proteins. Expression analysis revealed that the MYB-related gene family has a wide expression profile in maize and soybean development and plays important roles in development and stress responses. We hypothesize that MYB-related proteins initially diversified through three major expansions and domain shuffling, but remained relatively conserved throughout the subsequent plant evolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".