MétaCan
Menu
Back to cohort
Record W4317815831 · doi:10.1099/mgen.0.000908

The DataHarmonizer: a tool for faster data harmonization, validation, aggregation and analysis of pathogen genomics contextual information

2023· article· en· W4317815831 on OpenAlexafffundabout
Ivan S. Gill, Emma Griffiths, Damion Dooley, Rhiannon Cameron, Sarah Savić Kallesøe, Nithu Sara John, Anoosha Sehar, Gurinder Gosal, David Alexander, Madison Chapel, Matthew A. Croxen, Benjamin Delisle, Rachelle Di Tullio, Daniel Gaston, Ana T. Duggan, Jennifer L. Guthrie, Mark Horsman, Esha Joshi, Levon Kearny, Natalie Knox, Lynette Lau, Jason J. LeBlanc, Vincent Li, Pierre J. Lyons, Keith D. MacKenzie, Andrew G. McArthur, Emily M. Panousis, John Palmer, Natalie Prystajecky, Kerri Smith, Jennifer R. Tanner, Christopher Townend, Andrea D. Tyler, Gary Van Domselaar, William Hsiao

Bibliographic record

VenueMicrobial Genomics · 2023
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsSt. John’s Health Sciences CentreMcMaster UniversityUniversity of AlbertaSaskatchewan Disease Control LaboratoryPublic Health Agency of CanadaSimon Fraser UniversityInstitut National de Santé Publique du QuébecNova Scotia Health AuthorityBC Centre for Disease ControlHospital for Sick ChildrenPublic Health OntarioUniversity of British Columbia
FundersCanadian Institutes of Health ResearchMichael G. DeGroote Institute for Infectious Disease Research, McMaster UniversityGenome British ColumbiaMichael Smith Health Research BCGenome Canada
KeywordsMetadataData sharingComputer scienceData scienceHarmonizationInteroperabilityBig dataData integrationUsabilityWorld Wide WebDatabaseData miningMedicine

Abstract

fetched live from OpenAlex

Pathogen genomics is a critical tool for public health surveillance, infection control, outbreak investigations as well as research. In order to make use of pathogen genomics data, they must be interpreted using contextual data (metadata). Contextual data include sample metadata, laboratory methods, patient demographics, clinical outcomes and epidemiological information. However, the variability in how contextual information is captured by different authorities and how it is encoded in different databases poses challenges for data interpretation, integration and their use/re-use. The DataHarmonizer is a template-driven spreadsheet application for harmonizing, validating and transforming genomics contextual data into submission-ready formats for public or private repositories. The tool's web browser-based JavaScript environment enables validation and its offline functionality and local installation increases data security. The DataHarmonizer was developed to address the data sharing needs that arose during the COVID-19 pandemic, and was used by members of the Canadian COVID Genomics Network (CanCOGeN) to harmonize SARS-CoV-2 contextual data for national surveillance and for public repository submission. In order to support coordination of international surveillance efforts, we have partnered with the Public Health Alliance for Genomic Epidemiology to also provide a template conforming to its SARS-CoV-2 contextual data specification for use worldwide. Templates are also being developed for One Health and foodborne pathogens. Overall, the DataHarmonizer tool improves the effectiveness and fidelity of contextual data capture as well as its subsequent usability. Harmonization of contextual information across authorities, platforms and systems globally improves interoperability and reusability of data for concerted public health and research initiatives to fight the current pandemic and future public health emergencies. While initially developed for the COVID-19 pandemic, its expansion to other data management applications and pathogens is already underway.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.021
metaresearch head score (Gemma)0.040
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.025
Threshold uncertainty score0.111

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0210.040
Meta-epidemiology (narrow)0.0040.003
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0070.006
Science and technology studies0.0020.002
Scholarly communication0.0080.010
Open science0.0050.013
Research integrity0.0020.006
Insufficient payload (model declined to judge)0.0250.014

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.079
GPT teacher head0.313
Teacher spread0.234 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations19
Published2023
Admission routes3
Has abstractyes

Explore more

Same venueMicrobial GenomicsSame topicResearch Data Management PracticesFrench-language works237,207