MétaCan
Menu
Back to cohort
Record W6929813621 · doi:10.5281/zenodo.10783606

The Public Health Environmental Surveillance Database (PHESD) - Delatolla Lab v1.1.0

2021· dataset· en· W6929813621 on OpenAlexaffabout

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2021
Typedataset
Languageen
FieldMathematics
TopicNumerical methods for differential equations
Canadian institutionsUniversity of Ottawa
Fundersnot available
KeywordsPublic healthSoftware deploymentWastewaterPublic health surveillancePandemicHazardous wasteHealth data

Abstract

fetched live from OpenAlex

Description The repository is used to store Canadian wastewater and other environmental surveillance data using the Public Health and Environmental Surveillance Open Data Model (PHES-ODM). Background Wastewater testing and surveillance (WWS) has a long history as a public health tool, helping us monitor polio outbreaks and track antimicrobial resistance. With the COVID-19 pandemic and the emergence of the SARS-CoV-2 virus, a new opportunity has emerged to leverage WWS to inform prevention and control of the pandemic. People infected with the SARS-CoV-2 virus shed the virus in their stool, meaning that the feces in wastewater systems can be checked to detect outbreaks and monitor for variants even before people become symptomatic. With over 200 wastewater testing sites across Canada, and over 2000 testing sites in over 50 countries, WWS has proven itself to be an effective tool in the fight against the virus. However, despite the rapid growth in this field and the swift deployment of WWS programs nationally, there is very little data sharing due to the absence of a centralized data repository with controlled vocabulary. This absence has led to varied data and assay quality, inconsistent reporting of wastewater test results, little adjustment of wastewater results, and has slowed the development of wastewater-based epidemiology. That is why we developed the Public Health Environmental Surveillance Database (PHESD). This centralized database, funded by CoVaRR Net, the Canadian Institute for Health network, serves as an central depository for open access wastewater surveillance data. This will shrink the delay between the measurement and analysis, and provide more data for better modeling, better collaboration, and better tools in the fight against the COVID-19 pandemic. To ensure that the PHESD is usable for all users, we use the Public Health and Environmental Surveillance Open Data Model (PHES-ODM). The ODM an open, standard approach to share wastewater surveillance data. The ODM helps the wastewater surveillance community collaborate by allowing teams worldwide to share their data in a common structure. The model uses a relational dictionary to map more than 150 variables into 10 tables, and offers documentation on how to use the model, template files to record data in the ODM format, and scripts of code to set up a relational database according to the ODM schema. Current Objectives - Easily share wastewater data. The PHESD is focused on and built for Canadian data, but we are happy to revieve data from anywhere. The goal is to create one space to easily share data between researchers and programs. - Allow detailed wastewater data. This is a distinction between PHESD and other repositories - with the PHESD we allow for the storage of even detailed data. This means storing details on variants and sequencing, among other details. - Easy to add data. This database will use tPublic Health and Environmental Surveillance Open Data Model (PHES-ODM) to allow wastewater testing labs to share their using an open access, open science approach in a shared format. To help support labs and other data custodians is adopting the ODM format for the PHESD, we have developed a suite of tools to automate data validation using the Open Data Model validation schema and rule set. We can automatically add your data if it is in the ODM format on common data platforms like Dropbox, Sync, Google Drive, ArcGIS. This automatic data scraping can come from open access sources, but to preserve privacy for some labs we may scrap to a private version of this GitHub repository, which we then scrap onto this public-facing repository. Important Note: A method modification was applied on June 8th 2021 to the Ottawa wastewater data set. The magnitudes (not the shape of the curve) from June 8th to present have been retroactively modified on April 12, 2022 to better align the modified method to previous data.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.006
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Scholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.043
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.006
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0050.000
Scholarly communication0.0020.000
Open science0.0020.003
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0130.007

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.109
GPT teacher head0.330
Teacher spread0.221 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes2
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicNumerical methods for differential equationsFrench-language works237,207