MétaCan
Menu
← Back to cohort
Record W3145915389 · doi:10.1101/2021.03.30.434101

CanDIG: Secure Federated Genomic Queries and Analyses Across Jurisdictions

2021· preprint· en· W3145915389 on OpenAlexafffundabout
Lewis Jonathan Dursi, Zoltán Bozóky, Richard de Borja, Jimmy Li, David Bujold, Adam Lipski, Shaikh Farhan Rashid, Amanjeev Sethi, Neelam Memon, Dashaylan Naidoo, Felipe Coral-Sasso, Matthew L. Wong, P-O Quirion, Zhibin Lu, Samarth Agarwal, Kat Pavlov, Andrew Ponomarev, Mia Husić, Krista Pace, Samantha Palmer, Stephanie A. Grover, Sevan Hakgor, Lillian L. Siu, David Malkin, Carl Virtanen, Trevor J. Pugh, Pierre‐Étienne Jacques, Yann Joly, Steven J.M. Jones, Guillaume Bourque, Michael Brudno

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2021
Typepreprint
Languageen
FieldMedicine
TopicEthics in Clinical Research
Canadian institutionsMcGill Genome CentreSickKids FoundationUniversité de SherbrookeUniversity of TorontoHospital for Sick ChildrenMcGill UniversityOntario GenomicsCanada's Michael Smith Genome Sciences CentreInstitute of Cancer ResearchOntario Institute for Cancer ResearchProvidence Health CareProvincial Health Services AuthorityZymeworks (Canada)Princess Margaret Cancer CentreUniversity Health Network
FundersGarron Family Cancer CentreCHU Sainte-Justine FoundationFondation de l'Hôpital de Montréal pour enfantsBC Cancer FoundationChildhood Cancer CanadaHospital for Sick ChildrenTerry Fox Research InstituteAlberta Children's Hospital FoundationCanadian Institute for Advanced ResearchDalhousie UniversityBC Children's HospitalCanarieAlberta Cancer FoundationCanadian Institutes of Health ResearchJudith Jane Mason and Harold Stannett Williams Memorial FoundationChildren's Hospital FoundationCancerCare Manitoba FoundationMcMaster UniversitySickkids Research InstituteDalhousie Medical Research Foundation
KeywordsData sharingData scienceComputer scienceBig dataGenomicsData miningMedicineBiologyGenome

Abstract

fetched live from OpenAlex

Abstract Rapid expansions of bioinformatics and computational biology have broadened the collection and use of -omics data including genomic, transcriptomic, methylomic and a myriad of other health data types, in the clinic and the laboratory. Both clinical and research uses of such data require co-analysis with large datasets, for which participant privacy and the need for data custodian controls must remain paramount. This is particularly challenging in multi-jurisdictional settings, such as Canada, where health privacy and security requirements are often heterogeneous. Data federation presents a solution to this, allowing for integration and analysis of large datasets from various sites while abiding by local policies. The Canadian Distributed Infrastructure for Genomics platform (CanDIG) enables federated querying and analysis of -omics and health data while keeping that data local and under local control. It builds upon existing infrastructures to connect five health and research institutions across Canada, relies heavily on standards and tooling brought together by the Global Alliance for Genomics and Health (GA4GH), implements a clear division of responsibilities among its participants and adheres to international data sharing standards. Participating researchers and clinicians can therefore contribute to and quickly access a critical mass of -omics data across a national network in a manner that takes into account the multi-jurisdictional nature of our privacy and security policies. Through this, CanDIG gives medical and research communities the tools needed to use and analyze the ever-growing amount of -omics data available to them in order to improve our understanding and treatment of various conditions and diseases. CanDIG is being used to make genomic and phenotypic data available for querying across Canada as part of data sharing for five leading pan-Canadian projects including the Terry Fox Comprehensive Cancer Care Centre Consortium Network (TF4CN) and Terry Fox PRecision Oncology For Young peopLE (PROFYLE), and making data from provincial projects such as POG (Personalized Onco- Genomics) more widely available.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.017
metaresearch head score (Gemma)0.041
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesOpen science
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.992
Threshold uncertainty score0.231

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0170.041
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0030.004
Science and technology studies0.0040.005
Scholarly communication0.0100.008
Open science0.0080.026
Research integrity0.0040.003
Insufficient payload (model declined to judge)0.0080.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.175
GPT teacher head0.454
Teacher spread0.279 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designBench or experimental
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2021
Admission routes3
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)→Same topicEthics in Clinical Research→French-language works237,207→