EpiShare: an open platform to securely share epigenomic data
Notice bibliographique
Résumé
EpiShare is an open science project involving the International Human Epigenome Consortium (IHEC) and ENCODE, developing tools and APIs to increase accessibility of epigenomic data. It does so by using and contributing to standards established by the Global Alliance for Genomics and Health (GA4GH). Here, we present EpiShare’s recent initiatives and latest tools releases. First, we performed an in-depth study of the relevant legislation, ethical standards, and best practice documents to develop the Data Privacy Assessment Tool for Health (D-PATH). D-PATH is a unique tool to ensure that EpiShare’s data sharing activities meet the applicable ethical, legal, professional requirements. While the current version is taking into account the particular needs of EpiShare (physically located in Quebec while processing data from Canadian and international cohorts), D-PATH could be transposed to other data sharing projects, as we plan on expanding it to account for laws and policies from other jurisdictions. We also implemented an open-source clinical and phenotypical metadata service around Phenopackets, offering interoperability between other data exchange formats (FHIR, mCODE). This service promotes usage of existing biomedical ontologies and controlled vocabularies for annotations, and captures experiments metadata and their relationship to phenotypic descriptions. Consequently, we are participating in the GA4GH Clin/Pheno Work Stream to help develop the Phenopackets standard. Additionally, we are involved in the GA4GH REWS, contributing to the development of Data Access Committee Review Standards (DACReS). DACReS will identify areas of best practices for procedural standards to drive consistency and robust reviews for data access requests to genomic, epigenomics and health-related data. EpiShare will also play a leading role in the REWS Standard Genomic Data Licenses and Agreements project. Lastly, EpiShare has participated in the elaboration of the original rnaget API specification, and implemented it within the IHEC Data Portal to retrieve transcriptomic data in proposed data formats.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,031 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,004 | 0,003 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,006 | 0,009 |
| Science ouverte | 0,005 | 0,020 |
| Intégrité de la recherche | 0,002 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,037 | 0,024 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».