MétaCan
Menu
Back to cohort
Record W3195398961 · doi:10.1093/biosci/biab080

Data Proliferation, Reconciliation, and Synthesis in Viral Ecology

2021· article· en· W3195398961 on OpenAlexafffund
Rory Gibb, Gregory F. Albery, Daniel J. Becker, Liam Brierley, Ryan Connor, Tad Dallas, Evan A. Eskew, Maxwell J. Farrell, Angela L. Rasmussen, Sadie J. Ryan, Amy R. Sweeny, Colin J. Carlson, Timothée Poisot

Bibliographic record

VenueBioScience · 2021
Typearticle
Languageen
FieldMedicine
TopicZoonotic diseases and public health
Canadian institutionsUniversité de MontréalUniversity of SaskatchewanUniversity of Toronto
FundersMedical Research CouncilInstitut de Valorisation des DonnéesNational Science Foundation
KeywordsHuman viromeHost (biology)EcologyMacroecologyBiologyEvolutionary ecologyEvolutionary biologyData scienceComputer scienceMetagenomicsBiodiversityGenetics

Abstract

fetched live from OpenAlex

Abstract The fields of viral ecology and evolution are rapidly expanding, motivated in part by concerns around emerging zoonoses. One consequence is the proliferation of host–virus association data, which underpin viral macroecology and zoonotic risk prediction but remain fragmented across numerous data portals. In the present article, we propose that synthesis of host–virus data is a central challenge to characterize the global virome and develop foundational theory in viral ecology. To illustrate this, we build an open database of mammal host–virus associations that reconciles four published data sets. We show that this offers a substantially richer view of the known virome than any individual source data set but also that databases such as these risk becoming out of date as viral discovery accelerates. We argue for a shift in practice toward the development, incremental updating, and use of synthetic data sets in viral ecology, to improve replicability and facilitate work to predict the structure and dynamics of the global virome.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.179
metaresearch head score (Gemma)0.513
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: Theoretical or conceptual
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.179
Threshold uncertainty score0.945

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1790.513
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0140.014
Science and technology studies0.0030.011
Scholarly communication0.0180.024
Open science0.0060.015
Research integrity0.0030.007
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.061
GPT teacher head0.336
Teacher spread0.275 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designTheoretical or conceptual
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations39
Published2021
Admission routes2
Has abstractyes

Explore more

Same venueBioScienceSame topicZoonotic diseases and public healthFrench-language works237,207