Evolutionary fingerprinting of protein-coding genes in RNA viruses
Bibliographic record
Abstract
Abstract RNA viruses evolve rapidly to adapt to changing host environments. Much of this adaptation occurs at the level of protein-coding genes. Thus, some of the most well-characterized examples of rapid adaptation have been found in virus proteins that are exposed on the surface of the viral particle, where they mediate host receptor binding and cell entry. To investigate whether surface-exposed proteins and other proteins encoded by viruses exhibit different patterns of evolution under selection, we analyzed 244 protein-coding genes from 28 species of RNA viruses representing 15 taxonomic families. First, we show that gene-wide rates of non-synonymous ( dN ) and synonymous ( dS ) substitutions do not differentiate between categories of proteins. To provide a more detailed comparison between genes, we inferred for each alignment the bivariate posterior distribution over a fixed grid of codon site-specific dN and dS values. This distribution is the gene’s ‘evolutionary fingerprint’. Next, we computed the Wasserstein distance for every pair of fingerprints, which is analogous to amount of work required to reshape one distribution to another. After compensating for differences in genetic variation among viruses and proteins, we found that surface-exposed proteins could not be distinguished from non-exposed proteins in the space induced by the Wasserstein distance matrix. However, surface-exposed proteins from enveloped viruses were significantly clustered apart from their counterparts in non-enveloped viruses. In contrast, there was no significant separation between these categories of viruses for proteins with polymerase activity. We show that this pattern is more consistent with relaxed purifying selection than adaptive evolution in proteins associated with viral envelopes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".