A human biomonitoring (HBM) Global Registry Framework: Further advancement of HBM research following the FAIR principles
Bibliographic record
Abstract
Data generated by the rapidly evolving human biomonitoring (HBM) programmes are providing invaluable opportunities to support and advance regulatory risk assessment and management of chemicals in occupational and environmental health domains. However, heterogeneity across studies, in terms of design, terminology, biomarker nomenclature, and data formats, limits our capacity to compare and integrate data sets retrospectively (reuse). Registration of HBM studies is common for clinical trials; however, the study designs and resulting data collections cannot be traced easily. We argue that an HBM Global Registry Framework (HBM GRF) could be the solution to several of challenges hampering the (re)use of HBM (meta)data. The aim is to develop a global, host-independent HBM registry framework based on the use of harmonised open-access protocol templates from designing, undertaking of an HBM study to the use and possible reuse of the resulting HBM (meta)data. This framework should apply FAIR (Findable, Accessible, Interoperable and Reusable) principles as a core data management strategy to enable the (re)use of HBM (meta)data to its full potential through the data value chain. Moreover, we believe that implementation of FAIR principles is a fundamental enabler for digital transformation within environmental health. The HBM GRF would encompass internationally harmonised and agreed open access templates for HBM study protocols, structured web-based functionalities to deposit, find, and access harmonised protocols of HBM studies. Registration of HBM studies using the HBM GRF is anticipated to increase FAIRness of the resulting (meta)data. It is also considered that harmonisation of existing data sets could be performed retrospectively. As a consequence, data wrangling activities to make data ready for analysis will be minimised. In addition, this framework would enable the HBM (inter)national community to trace new HBM studies already in the planning phase and their results once finalised. The HBM GRF could also serve as a platform enhancing communication between scientists, risk assessors, and risk managers/policy makers. The planned European Partnership for the Assessment of Risk from Chemicals (PARC) work along these lines, based on the experience obtained in previous joint European initiatives. Therefore, PARC could very well bring a first demonstration of first essential functionalities within the development of the HBM GRF.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.465 | 0.330 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.004 | 0.006 |
| Bibliometrics | 0.014 | 0.015 |
| Science and technology studies | 0.005 | 0.016 |
| Scholarly communication | 0.029 | 0.049 |
| Open science | 0.013 | 0.028 |
| Research integrity | 0.010 | 0.013 |
| Insufficient payload (model declined to judge) | 0.007 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".