Author response: Precise assembly of complex beta sheet topologies from de novo designed building blocks
Bibliographic record
Abstract
A protein is made up of a sequence of amino acids and must fold into a specific three-dimensional structure if it is to work correctly. The structure is formed by segments of the protein adopting specific shapes, the two most common shapes being alpha helices and beta strands. Beta strands commonly interact with each other to form regions called beta sheets. Researchers trying to design proteins with new abilities have managed to create proteins that contain up to five beta strands and four alpha helices. Larger and more complex proteins are more challenging to make because there are many different ways that a protein can fold. It is also difficult to understand how complex structures such as large beta sheets emerged naturally, over the course of evolution. King et al. have now used computer modeling to explore how a large, complex beta sheet might form. In the model, one small, newly designed protein was inserted into another so that their beta sheets merged to form a single extended sheet. The model then stabilized this structure by changing the amino acids found at the points where the two proteins met. King et al. were then able to synthesize these new proteins in bacteria and use a technique called X-ray crystallography to determine the structure of two of them. The structures closely matched the computer models; one protein contained a six-stranded beta sheet, and the other had a seven-stranded beta sheet. The folds of the two designed proteins were then compared with those found in a database that classifies proteins on the basis of their structure. The beta sheets in the designed proteins did not match the protein structures in the database, which suggests that the designed proteins contained new types of folds. In the future, the technique used by King et al. could be used to design other large and complex beta sheet structures. Furthermore, the results suggest that such large structures could have evolved naturally through the combination of smaller, less complex proteins.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.006 | 0.004 |
| Insufficient payload (model declined to judge) | 0.225 | 0.105 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".