Sequence specificity analysis of the SETD2 protein lysine methyltransferase and discovery of a SETD2 super-substrate
Bibliographic record
Abstract
SETD2 catalyzes methylation at lysine 36 of histone H3 and it has many disease connections. We investigated the substrate sequence specificity of SETD2 and identified nine additional peptide and one protein (FBN1) substrates. Our data showed that SETD2 strongly prefers amino acids different from those in the H3K36 sequence at several positions of its specificity profile. Based on this, we designed an optimized super-substrate containing four amino acid exchanges and show by quantitative methylation assays with SETD2 that the super-substrate peptide is methylated about 290-fold more efficiently than the H3K36 peptide. Protein methylation studies confirmed very strong SETD2 methylation of the super-substrate in vitro and in cells. We solved the structure of SETD2 with bound super-substrate peptide containing a target lysine to methionine mutation, which revealed better interactions involving three of the substituted residues. Our data illustrate that substrate sequence design can strongly increase the activity of protein lysine methyltransferases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".