A comparative study of convicilin storage protein gene sequences in species of the tribe Vicieae
Bibliographic record
Abstract
Convicilins, a set of seed storage proteins, differ from vicilins, a related group of seed storage proteins, mainly because of the presence of the N-terminal extension, an additional sequence of amino acids in the sequence corresponding to the first exon. Convicilins have been described only in species of the legume tribe Vicieae. One or two genes for convicilins have been identified in most species of this tribe. The genus Pisum is the main exception, since two genes have been identified in most of its species. Thirty-four new convicilin gene sequences from 29 different species (Lathyrus, Lens, Pisum, and Vicia spp.) have been analyzed here. Convicilin gene sequences are generally organized in 6 exons, but in some instances one of the internal introns (2nd or 4th) is lost. In these 29 species, the N-terminal extension is formed by a stretch of 99 to 196 amino acids particularly rich in polar and charged amino acids (on average, it contains 29.43% glutamic acid and 15.38% arginine residues). This N-terminal extension has the characteristics of an intrinsically unstructured region (IUR), one of the categories of protein "degenerate sequences". A comparative analysis indicates that the N-terminal extension sequence has evolved faster than the surrounding sequence, which is common to all vicilins, and it evolved mainly through a series of duplications of short internal sequences and triplet expansions, the predominant one being GAA. This agrees with the evolution of IURs, which is faster than the evolution of surrounding sequences and is mainly due to replication slippage and unequal crossover recombination. Alternative maximum-likelihood trees of phylogenetic relationships among the 29 Vicieae species based on the convicilin exon sequences are presented and discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".