Circular permutation of concanavalin A: why the rarest protein modification in nature came to be
Bibliographic record
Abstract
Circular permutation is a concept that means rearranging the order of elements placed around a circle. This is relevant to biology because many proteins have their termini located close together in 3D space, thus mimicking a highly twisted yet closed loop. With regards to structure, where a protein begins and ends (the position of its N- and C-termini, respectively) may be less important for folding and function than the ordering of amino acids and domains around the loop. Thus, you can sometimes produce almost the same protein structure from two very differently ordered linear protein sequences (see Figure A). Such protein circular permutations have in fact arisen hundreds of time in nature, where protein domains are reordered at the DNA level but maintain their circular order and 3D structure at the protein level. Intriguingly, there is but one known example in nature of a protein undergoing circular permutation after it has been translated – and that is the story of concanavalin A (conA). ConA is a lectin (a carbohydrate-binding protein) that accumulates to high levels in seeds of the jack bean (Canavalia ensiformis). Although conA is widely used in carbohydrate affinity chromatography to separate glycoproteins and polysaccharides, its actual biological function in jack bean is unclear. Following translation, pro-conA undergoes a unique series of peptide cleavages and a transpeptidation reaction, resulting in a circular permutation of the mature conA (Figure B). Despite a detailed understanding of mature conA structure and its biophysical properties, the reasons for the evolution of such unusual post-translational modifications are unknown. In this issue, Nonis et al. (2021b) address this question by characterizing the 3D structure and biophysical properties of pro-conA in comparison to conA. Solving the pro-conA crystal structure revealed that its individual subunit (tertiary) structure was very similar to conA, including the folding of the functionally important carbohydrate-binding domains and metal-ligand-binding sites. And the carbohydrate-binding affinities of pro-conA relative to conA didn’t change when tested with mannose. However, conA typically exists as a dimer of dimers, and there were structural differences at the intermolecular interfaces of pro-conA that appeared to affect its dimer-dimer interactions. Indeed, analytical ultracentrifugation experiments revealed that the well-characterized pH-dependent shift in the dimer-tetramer equilibrium of conA was absent in pro-conA. Circular dichroism analysis then demonstrated that pro-conA is less stable than conA at low pH and at high temperature. These structural findings all indicated the cleavage and circular permutation of pro-conA provides a functional benefit in the form of greater protein stability. Questions remained about how the complicated post-translational processing of pro-conA could have evolved. Using an in vitro reaction assay, the authors demonstrated that a single asparagine endopeptidase, CeAEP1, is capable of catalyzing all the peptide cleavage and transpeptidation reactions that convert pro-conA to conA. Previous studies had demonstrated that CeAEP1 favors cleavage rather than transpeptidation for substrates other than conA (Bernath-Levin et al., 2015); therefore, the improved transpeptidation efficiency toward pro-conA shows that catalytic preference of CeAEP1 is substrate-dependent. Upon solving the crystal structure for CeAEP1 and comparing it to other structural models for asparaginyl endopeptidases (Haywood et al., 2018), the authors were able to provide new insights into the enzymatic basis for CeAEP1’s specific transpeptidation activity towards conA. Understanding the biochemistry behind the transpeptidation activities of asparagine endopeptidases greatly facilitates their use as valuable tools to modify polypeptide structure in the fields of protein engineering and synthetic biology (Nonis et al., 2021a). Circular permutation of conA. A, Circular permutation of proteins involves considering protein structure as a loop, then relocating the N and C termini within the polypeptide sequence without altering the sequential arrangement of folding domains. B, The post-translational processing of pro-conA by CeAEP1 Adapted from Nonis et al. (2021b), Figure 1. Circular permutation of conA. A, Circular permutation of proteins involves considering protein structure as a loop, then relocating the N and C termini within the polypeptide sequence without altering the sequential arrangement of folding domains. B, The post-translational processing of pro-conA by CeAEP1 Adapted from Nonis et al. (2021b), Figure 1.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.013 | 0.013 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".