Positive and Negative Cognate Amino Acid Bias Affects Compositions of Aminoacyl-tRNA Synthetases and Reflects Functional Constraints on Protein Structure
Bibliographic record
Abstract
By comparing phylogenetically related tRNA synthetases (enzymes that specifically aminoacylate tRNAs), a controlled natural experiment can reveal synthetases' coevolution with their cognate amino acid substrate.Analyses of metabolic cost minimization confirm the existence of cognate avoidance in tRNA synthetase compositions.It is found that cognate avoidance increases and decreases, respectively, with proteomic amino acid usage in Escherichia coli and Bacillus subtilis.In E. coli, cognate avoidance did not decrease with cognate abundance in the colon, but decrease with tRNA synthetase editing sites and cognate impact on protein structure, revealing that hydrophobic interactions and beta-sheet formation constrain the folding of tRNA synthetase classes I and II, respectively.Analyses of cognate bias yield information on how proteins function in E. coli, because function constrains cognate bias.In B. subtilis, cognate avoidance occurred for rare and abundant amino acids in the soil, positive bias existed for cognates with intermediate abundances.Presumably, life history strategies (endosymbiont versus free living) and environmental compositions modulate cost minimization of amino acid usages.Avoidance of 'expensive' residues in tRNA synthetases is inversely proportional to cognate avoidance and protein size in E. coli, but not B. subtilis.In relation to cost minimization, E. coli's predictable environment perhaps enabled to reach evolutionary (balancing) equilibrium between various factors affecting protein synthesis costs, where decreasing costs through one factor requires increasing other costs.Phylogenetically controlled comparisons can detect more statistically significant biases than similar analyses using randomly selected control proteins, stressing the power of carefully designed natural experiments.This work confirms the importance of biosynthetic cost minimization, and challenges neutralistic approaches of biomolecular evolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".