A Minimalist Model for Exploring Conformational Effects on the Electrospray Charge State Distribution of Proteins
Bibliographic record
Abstract
The electrospray ionization (ESI) charge state distribution of proteins is highly sensitive to the protein structure in solution. Unfolded conformations generally form higher charge states than tightly folded structures. The current study employs a minimalist molecular dynamics model for simulating the final stages of the ESI process in order to gain insights into the physical reasons underlying this empirical relationship. The protein is described as a string of 27 beads ("residues"), 9 of which are negatively charged and represent possible protonation sites. The unfolded state of this bead string is a random coil, whereas the native conformation adopts a compact fold. The ESI process is simulated by placing the protein inside a solvent droplet with a 2.5 nm radius consisting of 1600 Lennard-Jones particles. In addition, the droplet contains 14 protons which are modeled as highly mobile point charges. Disintegration of the droplet rapidly releases the protein into the gas phase, resulting in average charge states of 4.8+ and 7.4+ for the folded and unfolded conformation, respectively. The protonation probabilities of individual residues in the folded state reveal a characteristic pattern, with values ranging from 0.2 to 0.8. In contrast, the protonation probabilities of the unfolded protein are more uniform and cover the range from 0.8 to 1.0. The origin of these differences can be traced back to a combination of steric and electrostatic effects. Residues exhibiting a small accessible surface area are less likely to capture a proton, an effect that is exacerbated by partial electrostatic shielding from nearby positive residues. Conversely, sites that are sterically exposed are associated with electrostatic funnels that greatly increase the likelihood of protonation. Unfolding enhances the steric and electrostatic exposure of protonation sites, thereby causing the protein to capture a greater number of protons during the droplet disintegration process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".