MétaCan
Menu
Back to cohort

The Dataharmonizer: a Tool for Faster Data Harmonization, Validation, Aggregation, and Analysis of Pathogen Genomics Contextual Information

2022· preprint· en· W4283583762 on OpenAlexaffabout
Ivan Gill, Emma Griffiths, Damion Dooley, Rhiannon Cameron, Sarah Savić Kallesøe, Nithu Sara John, Anoosha Sehar, Gurinder Gosal, David Alexander, Madison Chapel, Matthew A. Croxen, Benjamin Delisle, Rachelle Di Tullio, Daniel Gaston, Ana T. Duggan, Jennifer L. Guthrie, Mark Horsman, Esha Joshi, Levon Kearney, Natalie Knox, Lynette Lau, Jason J. LeBlanc, Vincent Li, Pierre J. Lyons, Keith D. MacKenzie, Andrew G. McArthur, Emilie Panousis, John Palmer, Natalie Prystajecky, Kerri Smith, Jennifer E. Tanner, Christopher Townend, Andrea Tyler, Gary Van Domselaar, William Hsiao

Bibliographic record

VenuePreprints.org · 2022
Typepreprint
Languageen
FieldDecision Sciences
TopicScientific Computing and Data Management
Canadian institutionsSt. John’s Health Sciences CentreBC Centre for Disease ControlMcMaster UniversityOttawa Public HealthNova Scotia Health AuthoritySimon Fraser UniversityPublic Health Agency of CanadaHospital for Sick ChildrenUniversity of British ColumbiaPublic Health OntarioUniversity of AlbertaSaskatchewan Disease Control LaboratoryInstitut National de Santé Publique du Québec
Fundersnot available
KeywordsMetadataData sharingData scienceHarmonizationComputer scienceInteroperabilityBig dataData integrationContextual designWorld Wide WebDatabaseData miningMedicine

Abstract

fetched live from OpenAlex

Pathogen genomics is a critical tool for public health surveillance, infection control, outbreak investigations, as well as research. In order to make use of pathogen genomics data, it must be interpreted using contextual data (metadata). Contextual data includes sample metadata, laboratory methods, patient demographics, clinical outcomes, and epidemiological information. However, the variability in how contextual information is captured by different authorities and how it is encoded in different databases poses challenges for data interpretation, integration, and its use/re-use. The DataHarmonizer is a template-driven spreadsheet application for harmonizing, validating, and transforming genomics contextual data into submission-ready formats for public or private repositories. The tool’s web browser-based JavaScript environment enables validation and its offline functionality and local installation increases data security. The DataHarmonizer was developed to address the data sharing needs that arose during the COVID-19 pandemic, and was used by members of the Canadian COVID Genomics Network (CanCOGeN) to harmonize SARS-CoV-2 contextual data for national surveillance and for public repository submission.In order to support coordination of international surveillance efforts, we have partnered with the Public Health Alliance for Genomic Epidemiology to also provide a template conforming to its SARS-CoV-2 contextual data specification for use worldwide. Templates are also being developed for One Health and foodborne pathogens. Overall, the DataHarmonizer tool improves the effectiveness and fidelity of contextual data capture as well as its subsequent usability. Harmonization of contextual information across authorities, platforms and systems globally improves interoperability and reusability of data for concerted public health and research initiatives to fight the current pandemic and future public health emergencies. While initially developed for the COVID-19 pandemic, its expansion to other data management applications and pathogens is already underway.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.023
metaresearch head score (Gemma)0.041
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.032
Threshold uncertainty score0.119

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0230.041
Meta-epidemiology (narrow)0.0040.003
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0090.006
Science and technology studies0.0020.002
Scholarly communication0.0080.011
Open science0.0050.013
Research integrity0.0020.006
Insufficient payload (model declined to judge)0.0320.016

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.321
GPT teacher head0.415
Teacher spread0.094 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2022
Admission routes2
Has abstractyes

Explore more

Same venuePreprints.orgSame topicScientific Computing and Data ManagementFrench-language works237,207