MétaCan
Menu
Back to cohort
Record W4406774024 · doi:10.2196/63906

Toward a Domain-Overarching Metadata Schema for Making Health Research Studies FAIR (Findable, Accessible, Interoperable, and Reusable): Development of the NFDI4Health Metadata Schema

2025· article· en· W4406774024 on OpenAlexvenueno aff
Haitham Abaza, Aliaksandra Shutsko, Sophie Anne Inès Klopfenstein, Carina Nina Vorisek, Carsten Oliver Schmidt, Claudia Brünings-Kuppe, Vera Clemens, Johannes Darms, Sabine Hanß, Timm Intemann, Franziska Jannasch, Elisa Kasbohm, Birte Lindstädt, Matthias Löbe, Katharina Nimptsch, Ute Nöthlings, Marisabel Gonzalez Ocanto, Tracy Bonsu Osei, Ines Perrar, Manuela Peters, Tobias Pischon, Ulrich Sax, Matthias B. Schulze, Florian Schwarz, Carolina Schwedhelm, Sylvia Thun, Dagmar Waltemath, Atinkut Alamirrew Zeleke, Wolfgang G. Müller, Martin Golebiewski

Bibliographic record

VenueJMIR Medical Informatics · 2025
Typearticle
Languageen
FieldComputer Science
TopicResearch Data Management Practices
Canadian institutionsnot available
Fundersnot available
KeywordsMetadataPreprintComputer scienceSchema (genetic algorithms)Information retrievalWorld Wide WebData science

Abstract

fetched live from OpenAlex

BACKGROUND: Despite wide acceptance in medical research, implementation of the FAIR (findability, accessibility, interoperability, and reusability) principles in certain health domains and interoperability across data sources remain a challenge. While clinical trial registries collect metadata about clinical studies, numerous epidemiological and public health studies remain unregistered or lack detailed information about relevant study documents. Making valuable data from these studies available to the research community could improve our understanding of various diseases and their risk factors. The National Research Data Infrastructure for Personal Health Data (NFDI4Health) seeks to optimize data sharing among the clinical, epidemiological, and public health research communities while preserving privacy and ethical regulations. OBJECTIVE: We aimed to develop a tailored metadata schema (MDS) to support the standardized publication of health studies' metadata in NFDI4Health services and beyond. This study describes the development, structure, and implementation of this MDS designed to improve the FAIRness of metadata from clinical, epidemiological, and public health research while maintaining compatibility with metadata models of other resources to ease interoperability. METHODS: Based on the models of DataCite, ClinicalTrials.gov, and other data models and international standards, the first MDS version was developed by the NFDI4Health Task Force COVID-19. It was later extended in a modular fashion, combining generic and NFDI4Health use case-specific metadata items relevant to domains of nutritional epidemiology, chronic diseases, and record linkage. Mappings to schemas of clinical trial registries and international and local initiatives were performed to enable interfacing with external resources. The MDS is represented in Microsoft Excel spreadsheets. A transformation into an improved and interactive machine-readable format was completed using the ART-DECOR (Advanced Requirement Tooling-Data Elements, Codes, OIDs, and Rules) tool to facilitate editing, maintenance, and versioning. RESULTS: The MDS is implemented in NFDI4Health services (eg, the German Central Health Study Hub and the Local Data Hub) to structure and exchange study-related metadata. Its current version (3.3) comprises 220 metadata items in 5 modules. The core and design modules cover generic metadata, including bibliographic information, study design details, and data access information. Domain-specific metadata are included in use case-specific modules, currently comprising nutritional epidemiology, chronic diseases, and record linkage. All modules incorporate mandatory, optional, and conditional items. Mappings to the schemas of clinical trial registries and other resources enable integrating their study metadata in the NFDI4Health services. The current MDS version is available in both Excel and ART-DECOR formats. CONCLUSIONS: With its implementation in the German Central Health Study Hub and the Local Data Hub, the MDS improves the FAIRness of data from clinical, epidemiological, and public health research. Due to its generic nature and interoperability through mappings to other schemas, it is transferable to services from adjacent domains, making it useful for a broader user community.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.073
metaresearch head score (Gemma)0.079
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesnone
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.995
Threshold uncertainty score0.387

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0730.079
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0100.009
Science and technology studies0.0030.004
Scholarly communication0.0100.014
Open science0.0050.012
Research integrity0.0030.006
Insufficient payload (model declined to judge)0.0030.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.383
GPT teacher head0.531
Teacher spread0.148 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainReproducibility
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations7
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Medical InformaticsSame topicResearch Data Management PracticesFrench-language works237,207