Tertiary structural motif sequence statistics enable facile prediction and design of peptides that bind anti-apoptotic Bfl-1 and Mcl-1
Notice bibliographique
Résumé
Abstract Understanding the relationship between protein sequence and structure well enough to rationally design novel proteins or protein complexes is a longstanding goal in protein science. The Protein Data Bank (PDB) is a key resource for defining sequence-structure relationships that has supported the development of critical resources such as rotamer libraries and backbone torsional statistics that quantify the probabilities of protein sequences adopting different structures. Here, we show that well-defined, non-contiguous structural motifs (TERMs) in the PDB can also provide rich information useful for protein-peptide interaction prediction and design. Specifically, we show that it is possible to rapidly predict the binding energies of peptides to Bcl-2 family proteins as accurately as can be done with widely used structure-based tools, without explicit atomistic modeling. One benefit of a TERM-based approach is that prediction performance is less sensitive to the details of the input structure than are methods that evaluate energies using precise atomic coordinates. We show that protein design using TERM energies (dTERMen) can generate highly novel and diverse peptides to target anti-apoptotic proteins Bfl-1 and Mcl-1. 15 of 17 peptides designed using dTERMen bound tightly to their intended targets, and these peptides have just 15 - 38% sequence identity to any known native Bcl-2 family protein ligand. High-resolution structures of four designed peptides bound to their targets provided opportunities to analyze strengths and limitations of this approach. Dramatic success designing peptides using dTERMen, which comprised going from input structure to experimental validation of high-affinity binders in approximately one month, provides strong motivation for further developing TERM-based approaches to design.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,001 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».