Reprogramming Caspase-7 Specificity by Regio-Specific Mutations and Selection Provides Alternate Solutions for Substrate Recognition
Bibliographic record
Abstract
The ability to routinely engineer protease specificity can allow us to better understand and modulate their biology for expanded therapeutic and industrial applications. Here, we report a new approach based on a caged green fluorescent protein (CA-GFP) reporter that allows for flow-cytometry-based selection in bacteria or other cell types enabling selection of intracellular protease specificity, regardless of the compositional complexity of the protease. Here, we apply this approach to introduce the specificity of caspase-6 into caspase-7, an intracellular cysteine protease important in cellular remodeling and cell death. We found that substitution of substrate-contacting residues from caspase-6 into caspase-7 was ineffective, yielding an inactive enzyme, whereas saturation mutagenesis at these positions and selection by directed evolution produced active caspases. The process produced a number of nonobvious mutations that enabled conversion of the caspase-7 specificity to match caspase-6. The structures of the evolved-specificity caspase-7 (esCasp-7) revealed alternate binding modes for the substrate, including reorganization of an active site loop. Profiling the entire human proteome of esCasp-7 by N-terminomics demonstrated that the global specificity toward natural protein substrates is remarkably similar to that of caspase-6. Because the esCasp-7 maintained the core of caspase-7, we were able to identify a caspase-6 substrate, lamin C, that we predict relies on an exosite for substrate recognition. These reprogrammed proteases may be the first tool built with the express intent of distinguishing exosite dependent or independent substrates. This approach to specificity reprogramming should also be generalizable across a wide range of proteases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".