The role of local structure propensity and nonnative interactions in protein folding
Bibliographic record
Abstract
Despite many years of research, the roles of local as well as nonnative interactions in protein folding remain poorly understood. To bridge this gap in knowledge, this thesis is an investigation of the contributions of local structure propensity as well as nonnative interactions to the energetics of protein folding. To address the energetic principles of local structure propensity, a multiple substitution strategy was employed where the effect on protein folding kinetics of a large subset of mutants at a single surface-exposed position is examined. By taking advantage of the characteristic feature of folding transition states in being less tightly packed than the folded state, I demonstrated that folding kinetics may generally provide the best means to characterize the energetics of local structure propensities with less influence of packing, as opposed to equilibrium stability studies. Application of this principle to two surface exposed β strand positions in the Fyn SH3 domain suggested that it is possible to obtain a context-independent assessment of local structure propensities through transition state analysis. Once applied to a 310 helix position in the domain, an experimental propensity scale for 310 helices was established for the first time. Therefore, with the rapid growth in structural databases, the multiple substitution approach provides a means to characterize informational propensity values of a larger number of motifs, besides helices and sheets. My findings suggest that, during folding, normative interactions may mask the energetic contributions of local structure propensities. Moreover, quite contrary to the common perception, I observed that normative interactions can speed up folding and that the propensity to form normative contacts is not uniform along protein sequence. Thus, extra caution needs to be exercised in interpreting the results of Φ value analysis and those of pure native centric models, which assume that nonnative interactions are either nonexistent or do not positively influence folding.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".