Author response: The kinetoplastid-infecting Bodo saltans virus (BsV), a window into the most abundant giant viruses in the sea
Bibliographic record
Abstract
In oceans, rivers and lakes, there are about a million viruses in every milliliter of water. Most of these viruses are tiny, often 10 or 100 times smaller than bacteria. However, a few reach a similar size and complexity to bacteria, and so stand out as relative giants. Relative to other viruses, Giant Viruses have much more DNA in their genome, which in turn provides the genetic template to produce the proteins that allow viruses to reproduce largely independently of its host. Typically, more than half of the genes encoded by Giant Viruses have no evident similarity to genes from other viruses or cellular life. Sequencing DNA from ocean water suggests that Giant Viruses are abundant and ecologically important; yet, few have been isolated from the microbes that they infect. Without being able to study Giant Viruses in the laboratory, little can be known about their biology, the way they infect their hosts, and their broader influence on aquatic life. Deeg et al. have now isolated and characterized the giant Bodo saltans virus (BsV), a Giant Virus that infects an ecologically important microbe commonly found in aquatic environments. Sequencing the genome of BsV revealed many previously unknown genes, as well as several unusual features. For example, the genome contains movable genetic elements that might help to fend off other giant viruses by cutting their genomes. In addition, the set of genes used by BsV to translate mRNA templates into proteins differs from those found in other giant viruses, implying that they are not derived from a more complex common ancestor. The size of the genome appears to have grown rapidly by the duplication of genes at the end of the genome – a feature known as a genomic accordion. The identity of the duplicated genes suggests that there is an evolutionary arms race with its host that forces genome expansion. Further studies of the BsV genome could help researchers to understand the origin of gigantism in the genomes of giant viruses.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.014 | 0.008 |
| Insufficient payload (model declined to judge) | 0.177 | 0.076 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".