Comparing the diversity and relative abundance of free and particle-associated aquatic viruses
Bibliographic record
Abstract
ABSTRACT Metagenomics has enabled rapid increases in virus discovery, in turn permitting revisions of viral taxonomy and our understanding of the ecology of viruses and their hosts. Inspired by recent discoveries of large viruses prevalent in the environment, we re-assessed the longstanding approach of filtering water through small pore-size filters to separate viruses from cells before sequencing. We studied assembled contigs derived from < 0.45 μm and > 0.45 μm size fractions that were annotated as viral to determine the diversity and relative abundances of virus groups from each fraction. Virus communities were vastly different when comparing the size fractions, indicating that analysis of either fraction alone would provide only a partial perspective of environmental viruses. At the level of virus order/family we observed highly diverse and distinct virus communities in the > 0.45 μm size fractions, whereas the < 0.45 μm size fractions were comprised primarily of highly diverse Caudovirales. The relative abundances of Caudovirales for which hosts could be inferred varied widely between size fractions with higher relative abundances of cyanophages in the > 0.45 μm size fractions potentially indicating replication within cells during ongoing infections. Many of the Mimiviridae and Phycodnaviridae , and all Iridoviridae and Poxviridae were detected exclusively in the often disregarded > 0.45 μm size fractions. In addition to observing unique virus communities associated with each size fraction, we detected viruses common to both fractions and argue that these are candidates for further exploration because they may be the product of ongoing or recent lytic events. IMPORTANCE Most studies of aquatic virus communities analyze DNA sequences derived from the smaller, “free virus” size fraction. Our study demonstrates that analysis of virus communities using only the smaller size fraction can lead to erroneously low diversity estimates for many of the larger viruses such as Mimiviridae, Phycodnaviridae, Iridoviridae , and Poxviridae , whereas analyzing only the larger, > 0.45 μm size fraction can lead to underestimates of Caudovirales diversity and relative abundance. Similarly, our data shows that examining only the smaller size fraction can lead to underestimation of virophage and cyanophage relative abundances that could, in turn, cause researchers to assume their limited ecological importance. Given the considerable differences we observed in this study, we recommend cautious interpretations of environmental virus community assemblages and dynamics when based on metagenomic data derived from different size fractions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".