EpiShare: an open platform to securely share epigenomic data
Bibliographic record
Abstract
EpiShare is an open science project involving the International Human Epigenome Consortium (IHEC) and ENCODE, developing tools and APIs to increase accessibility of epigenomic data. It does so by using and contributing to standards established by the Global Alliance for Genomics and Health (GA4GH). Here, we present EpiShare’s recent initiatives and latest tools releases. First, we performed an in-depth study of the relevant legislation, ethical standards, and best practice documents to develop the Data Privacy Assessment Tool for Health (D-PATH). D-PATH is a unique tool to ensure that EpiShare’s data sharing activities meet the applicable ethical, legal, professional requirements. While the current version is taking into account the particular needs of EpiShare (physically located in Quebec while processing data from Canadian and international cohorts), D-PATH could be transposed to other data sharing projects, as we plan on expanding it to account for laws and policies from other jurisdictions. We also implemented an open-source clinical and phenotypical metadata service around Phenopackets, offering interoperability between other data exchange formats (FHIR, mCODE). This service promotes usage of existing biomedical ontologies and controlled vocabularies for annotations, and captures experiments metadata and their relationship to phenotypic descriptions. Consequently, we are participating in the GA4GH Clin/Pheno Work Stream to help develop the Phenopackets standard. Additionally, we are involved in the GA4GH REWS, contributing to the development of Data Access Committee Review Standards (DACReS). DACReS will identify areas of best practices for procedural standards to drive consistency and robust reviews for data access requests to genomic, epigenomics and health-related data. EpiShare will also play a leading role in the REWS Standard Genomic Data Licenses and Agreements project. Lastly, EpiShare has participated in the elaboration of the original rnaget API specification, and implemented it within the IHEC Data Portal to retrieve transcriptomic data in proposed data formats.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.031 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.009 |
| Open science | 0.005 | 0.020 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.037 | 0.024 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".