Restricted dead‐end elimination: Protein redesign with a bounded number of residue mutations
Bibliographic record
Abstract
Dead-end elimination (DEE) has emerged as a powerful structure-based, conformational search technique enabling computational protein redesign. Given a protein with n mutable residues, the DEE criteria guide the search toward identifying the sequence of amino acids with the global minimum energy conformation (GMEC). This approach does not restrict the number of permitted mutations and allows the identified GMEC to differ from the original sequence in up to n residues. In practice, redesigns containing a large number of mutations are often problematic when taken into the wet-lab for creation via site-directed mutagenesis. The large number of point mutations required for the redesigns makes the process difficult, and increases the risk of major unpredicted and undesirable conformational changes. Preselecting a limited subset of mutable residues is not a satisfactory solution because it is unclear how to select this set before the search has been performed. Therefore, the ideal approach is what we define as the kappa-restricted redesign problem in which any kappa of the n residues are allowed to mutate. We introduce restricted dead-end elimination (rDEE) as a solution of choice to efficiently identify the GMEC of the restricted redesign (the kappaGMEC). Whereas existing approaches require n-choose-kappa individual runs to identify the kappaGMEC, the rDEE criteria can perform the redesign in a single search. We derive a number of extensions to rDEE and present a restricted form of the A* conformation search. We also demonstrate a 10-fold speed-up of rDEE over traditional DEE approaches on three different experimental systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".