Giant RNA genomes: Roles of host, translation elongation, genome architecture, and proteome in nidoviruses
Bibliographic record
Abstract
have the largest known RNA genomes of vertebrate and invertebrate viruses with 36.7 and 41.1 kb, respectively. The acquisition of a proofreading exoribonuclease (ExoN) by an ancestral nidovirus enabled crossing of the 20 kb barrier. Other factors constraining genome size variations in nidoviruses remain poorly defined. We assemble 76 genome sequences of invertebrate nidoviruses from >500.000 published transcriptome experiments and triple the number of known nidoviruses with >36 kb genomes, including a 64 kb RNA genome. Many of the identified viral lineages acquired putative enzymatic and other protein domains linked to genome size, host phyla, or virus families. The inserted domains may regulate viral replication and virion formation, or modulate infection otherwise. We classify ExoN-encoding nidoviruses into seven groups and four subgroups, according to canonical and noncanonical modes of viral replicase expression by ribosomes and genomic organization (reModes). The most-represented group employing the canonical reMode comprises invertebrate and vertebrate nidoviruses, including coronaviruses. Six groups with noncanonical reModes include invertebrate nidoviruses with 31-to-64 kb genomes. Among them are viruses with segmented genomes and viruses utilizing dual ribosomal frameshifting that we validate experimentally. Moreover, largest polyprotein length and genome size in nidoviruses show reMode- and host phylum-dependent relationships. We hypothesize that the polyprotein length increase in nidoviruses may be limited by the host-inherent translation fidelity, ultimately setting a nidovirus genome size limit. Thus, expansion of ExoN-encoding RNA virus genomes, the vertebrate/invertebrate host division, the control of viral replicase expression, and translation fidelity are interconnected.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".