BY-COVID D2.3 Enabling data discovery at source using beacon-like mechanisms
Bibliographic record
Abstract
Deliverable D2.3, titled "Enabling data discovery at source using beacon-like mechanisms," presents the outcomes and advancements achieved by Work Package 2 (WP2) partners within the BY-COVID project. The report delineates existing data discovery mechanisms and introduces novel cross-domain data discovery through extensions to the Beacon and the Beacon Network technologies, particularly focusing on its application to COVID-19 data. Throughout the duration of the project, significant progress has been made in expanding standard data discovery mechanisms. Collaborative efforts have resulted in the enhancement of Global Alliance for Genomics and Health (GA4GH) Beacon-based mechanisms, enabling efficient data discovery at its source. This deliverable describes eight different Beacon implementations from BY-COVID partners, and other institutions in Europe and other parts of the world (Canada and Australia). These beacons share cross-domain data, from viral genomes and epidemiology to rich patient information or combined viral and host genomes from the same donors. The achievements prove the commitment to enabling data discovery at its source. By extending Beacon technologies, the project has paved the way for enhanced data accessibility and interoperability across various research domains. Nevertheless, the deliverable shows that the use of popular models and dictionaries (like OMOP or ISARIC eCRF) is not enough to achieve solid interoperability, although it enormously reduces the harmonisation gap and opens the door to generation of tools to bridge these gaps. These advancements signify a crucial step forward in empowering researchers to efficiently access and use diverse datasets, ultimately contributing to more informed decision-making and research outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".