A comparative study of convicilin storage protein gene sequences in species of the tribe Vicieae
Bibliographic record
Abstract
Convicilins, a set of seed storage proteins, differ from vicilins, a related group of seed storage proteins, mainly because of the presence of the N-terminal extension, an additional sequence of amino acids in the sequence corresponding to the first exon. Convicilins have been described only in species of the legume tribe Vicieae. One or two genes for convicilins have been identified in most species of this tribe. The genus Pisum is the main exception, since two genes have been identified in most of its species. Thirty-four new convicilin gene sequences from 29 different species (Lathyrus, Lens, Pisum, and Vicia spp.) have been analyzed here. Convicilin gene sequences are generally organized in 6 exons, but in some instances one of the internal introns (2nd or 4th) is lost. In these 29 species, the N-terminal extension is formed by a stretch of 99 to 196 amino acids particularly rich in polar and charged amino acids (on average, it contains 29.43% glutamic acid and 15.38% arginine residues). This N-terminal extension has the characteristics of an intrinsically unstructured region (IUR), one of the categories of protein "degenerate sequences". A comparative analysis indicates that the N-terminal extension sequence has evolved faster than the surrounding sequence, which is common to all vicilins, and it evolved mainly through a series of duplications of short internal sequences and triplet expansions, the predominant one being GAA. This agrees with the evolution of IURs, which is faster than the evolution of surrounding sequences and is mainly due to replication slippage and unequal crossover recombination. Alternative maximum-likelihood trees of phylogenetic relationships among the 29 Vicieae species based on the convicilin exon sequences are presented and discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".