MétaCan
Menu
Back to cohort
Record W4403393007 · doi:10.1099/mgen.0.001293

The Canadian VirusSeq Data Portal and Duotang: open resources for SARS-CoV-2 viral sequences and genomic epidemiology

2024· article· en· W4403393007 on OpenAlexafffundabout
Erin E. Gill, Baofeng Jia, Carmen Lía Murall, Raphaël Poujol, Muhammad Zohaib Anwar, Nithu Sara John, Justin Richardsson, Ashley E. Hobb, Abayomi S. Olabode, Alexandru Lepsa, Ana T. Duggan, Andrea D. Tyler, Arnaud N’Guessan, Atul Kachru, Brandon Chan, Catherine Yoshida, Christina K. Yung, David Bujold, Dusan Andric, Edmund Su, Emma Griffiths, Gary Van Domselaar, Gordon Jolly, Heather Ward, Henrich Feher, Jared Baker, Jared T. Simpson, Jaser Uddin, Jiannis Ragoussis, Jon Eubank, Jörg H. Fritz, José Héctor Gálvez, Karen Fang, Kim Cullion, Leonardo Landa Rivera, Qian Xiang, Matthew A. Croxen, Mitchell Shiell, Natalie Prystajecky, Pierre-Olivier Quirion, Rosita Bajari, Samantha Rich, Samira Mubareka, Sandrine Moreira, Scott Cain, Steven G. Sutcliffe, Susanne A. Kraemer, Yelizar Alturmessov, Yann Joly, Marc Fiume, Terrance P. Snutch, Cindy Bell, Catalina López-Correa, Julie Hussin, Jeffrey B. Joy, Caroline Colijn, Paul M. K. Gordon, William Hsiao, Art F. Y. Poon, Natalie Knox, Mélanie Courtot, Lincoln Stein, Sarah P. Otto, Guillaume Bourque, B. Jesse Shapiro, Fiona S. L. Brinkman

Bibliographic record

VenueMicrobial Genomics · 2024
Typearticle
Languageen
FieldMedicine
TopicSARS-CoV-2 and COVID-19 Research
Canadian institutionsUniversity of CalgaryAIDS VancouverMila - Quebec Artificial Intelligence InstituteGenome CanadaMichael Smith Health Research BCSunnybrook Health Science CentreBC Centre for Disease ControlWomen and Children’s Health Research InstituteOntario GenomicsUniversity of AlbertaIndoc ResearchUniversité de MontréalMcGill Genome CentreUniversity of TorontoWestern UniversityMcGill UniversitySimon Fraser UniversityPublic Health Agency of CanadaOntario Institute for Cancer ResearchUniversity of British ColumbiaMontreal Heart Institute
FundersNational Cancer InstituteCanadian Institutes of Health ResearchInstitut de Valorisation des DonnéesFonds de Recherche du Québec - SantéMichael Smith Health Research BCAlliance de recherche numérique du CanadaGenome AlbertaPublic Health Agency of CanadaSimon Fraser UniversityGovernment of OntarioGenome CanadaNational Institutes of HealthOntario GenomicsGovernment of CanadaPublic Health AgencyInnovation, Science and Economic Development CanadaCanarie
KeywordsData sharingPandemicInteroperabilityBig dataSuiteGenomicsData accessComputer scienceData scienceWorld Wide WebGenomeCoronavirus disease 2019 (COVID-19)DatabaseBiologyPolitical scienceMedicineData miningGenetics

Abstract

fetched live from OpenAlex

The COVID-19 pandemic led to a large global effort to sequence SARS-CoV-2 genomes from patient samples to track viral evolution and inform the public health response. Millions of SARS-CoV-2 genome sequences have been deposited in global public repositories. The Canadian COVID-19 Genomics Network (CanCOGeN - VirusSeq), a consortium tasked with coordinating expanded sequencing of SARS-CoV-2 genomes across Canada early in the pandemic, created the Canadian VirusSeq Data Portal, with associated data pipelines and procedures, to support these efforts. The goal of VirusSeq was to allow open access to Canadian SARS-CoV-2 genomic sequences and enhanced, standardized contextual data that were unavailable in other repositories and that meet FAIR standards (Findable, Accessible, Interoperable and Reusable). In addition, the portal data submission pipeline contains data quality checking procedures and appropriate acknowledgement of data generators that encourages collaboration. From inception to execution, the portal was developed with a conscientious focus on strong data governance principles and practices. Extensive efforts ensured a commitment to Canadian privacy laws, data security standards, and organizational processes. This portal has been coupled with other resources, such as Viral AI, and was further leveraged by the Coronavirus Variants Rapid Response Network (CoVaRR-Net) to produce a suite of continually updated analytical tools and notebooks. Here we highlight this portal (https://virusseq-dataportal.ca/), including its contextual data not available elsewhere, and the Duotang (https://covarr-net.github.io/duotang/duotang.html), a web platform that presents key genomic epidemiology and modelling analyses on circulating and emerging SARS-CoV-2 variants in Canada. Duotang presents dynamic changes in variant composition of SARS-CoV-2 in Canada and by province, estimates variant growth, and displays complementary interactive visualizations, with a text overview of the current situation. The VirusSeq Data Portal and Duotang resources, alongside additional analyses and resources computed from the portal (COVID-MVP, CoVizu), are all open source and freely available. Together, they provide an updated picture of SARS-CoV-2 evolution to spur scientific discussions, inform public discourse, and support communication with and within public health authorities. They also serve as a framework for other jurisdictions interested in open, collaborative sequence data sharing and analyses.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.021
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesOpen science
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.996
Threshold uncertainty score0.435

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.021
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0070.012
Science and technology studies0.0040.001
Scholarly communication0.0070.005
Open science0.0040.008
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0750.035

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.130
GPT teacher head0.396
Teacher spread0.265 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2024
Admission routes3
Has abstractyes

Explore more

Same venueMicrobial GenomicsSame topicSARS-CoV-2 and COVID-19 ResearchFrench-language works237,207