Missense mutations: Backbone structure positional effects
Bibliographic record
Abstract
ABSTRACT Human diversity often manifests through single nucleotide polymorphisms (SNPs). Among these, missense mutations, or SNPs that alter amino acids, can modify a protein’s three-dimensional (3D) structure. This impacts its function and can potentially elicit diseases or affect drug interactions. Thus, understanding protein single point mutations is crucial for precision medicine, as it helps tailor treatments based on individual genetic variations. As atomic locations can be susceptible to any number of changes that might or might not affect function, we focus on the secondary structure to provide concrete results on possible protein structural deformation that may occur from missense mutations. We assess state-of-the-art structure prediction methods regarding backbone deformations caused by missense mutations. We categorize these deformations as local, distant , or global based on the proximity of structural changes to the mutation site. Our analysis utilizes a diverse dataset from the Protein Data Bank, comprising over 500 protein clusters with experimentally determined structures and documented mutations. Our findings indicate that missense mutations can significantly affect the accuracy of structure prediction methods. These mutations often lead to predicted structural changes even when the actual secondary structures remain unchanged, suggesting that current methods overestimate the impact of missense mutations. This issue is particularly evident in advanced prediction algorithms, which struggle to accurately model proteins with stable mutations. We also found that the addition of low-performing prediction methods during structural analysis can positively impact the results on some proteins, particularly those with low homology. Furthermore, proteins that form complexes or bind ligands—such as membrane and transport proteins—are inaccurately predicted due to the absence of extra-molecular interaction data in the models, highlighting how missense mutations can complicate accurate structure prediction. All code and data are available at https://github.com/ivanpmartell/pdb-sam .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".