Active site specificity profiling of the matrix metalloproteinase family: Proteomic identification of 4300 cleavage sites by nine MMPs explored with structural and synthetic peptide cleavage analyses
Bibliographic record
Abstract
Secreted and membrane tethered matrix metalloproteinases (MMPs) are key homeostatic proteases regulating the extracellular signaling and structural matrix environment of cells and tissues. For drug targeting of proteases, selectivity for individual molecules is highly desired and can be met by high yield active site specificity profiling. Using the high throughput Proteomic Identification of protease Cleavage Sites (PICS) method to simultaneously profile both the prime and non-prime sides of the cleavage sites of nine human MMPs, we identified more than 4300 cleavages from P6 to P6' in biologically diverse human peptide libraries. MMP specificity and kinetic efficiency were mainly guided by aliphatic and aromatic residues in P1' (with a ~32-93% preference for leucine depending on the MMP), and basic and small residues in P2' and P3', respectively. A wide differential preference for the hallmark P3 proline was found between MMPs ranging from 15 to 46%, yet when combined in the same peptide with the universally preferred P1' leucine, an unexpected negative cooperativity emerged. This was not observed in previous studies, probably due to the paucity of approaches that profile both the prime and non-prime sides together, and the masking of subsite cooperativity effects by global heat maps and iceLogos. These caveats make it critical to check for these biologically highly important effects by fixing all 20 amino acids one-by-one in the respective subsites and thorough assessing of the inferred specificity logo changes. Indeed an analysis of bona fide MEROPS physiological substrate cleavage data revealed that of the 37 natural substrates with either a P3-Pro or a P1'-Leu only 5 shared both features, confirming the PICS data. Upon probing with several new quenched-fluorescent peptides, rationally designed on our specificity data, the negative cooperativity was explained by reduced non-prime side flexibility constraining accommodation of the rigidifying P3 proline with leucine locked in S1'. Similar negative cooperativity between P3 proline and the novel preference for asparagine in P1 cements our conclusion that non-prime side flexibility greatly impacts MMP binding affinity and cleavage efficiency. Thus, unexpected sequence cooperativity consequences were revealed by PICS that uniquely encompasses both the non-prime and prime sides flanking the proteomic-pinpointed scissile bond.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".