Local-structural diversity and protein folding: Application to all-β off-lattice protein models
Bibliographic record
Abstract
Global measures of structural diversity within a distribution of biopolymers, such as the radius of gyration and percent native contacts, have proven useful in the analysis of simulation data for protein folding. In this paper we describe a statistical-based methodology to quantify the local structural variability of a distribution of biopolymers, applied to 46- and 69-"residue" off-lattice, three-color model proteins. Each folds into beta-barrel structures. First we perform a principal component analysis of all interbead distance variables for a large number of independent, converged Boltzmann-distributed samples of conformations collected at each of a wide range of temperatures. Next, the principal component vectors are subjected to orthogonal (varimax) rotation. The results are displayed on so-called "squared-loading" plots. These provide a quantitative measure of the contribution to the sample variance of the position of each residue relative to the others. Dominant structural elements, those having the largest structural diversity within the sampled distribution, are responsible for peaks and shoulders observed in the specific heat versus temperature curves, generated using the weighted histogram analysis method. The loading plots indicate that the local-structural diversity of these systems changes gradually with temperature through the folding transition but radically changes near the collapse transition temperature. The analysis of the structural overlap order statistic suggests that the 46-mer thermodynamic folding transition involves the native state and at least three other nearly native intermediates. In the case of the 46-mer protein model, data are generated at sufficiently low temperatures that squared-loading plots, coupled with cluster analysis, provide a local and energetic description of its glassy state.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".