Bibliographic record
Abstract
The TRIPLE Training Toolkit is part of the work performed by Work Package 6 (WP6) under Task 6.3 in the TRIPLE Project (Transforming Research through Linked Interdisciplinary Exploration). The project is funded by the European Commission, under Grant Agreement No. 863420 and will run for 42 months starting from October 2019. In light of the need for a common understanding of European Open Science advancements and to support the uptake of Open Science practices within SSH research and training communities, Task 6.3 produced two kinds of outputs. The TRIPLE Open Science Training Series is a series of 12 open and reusable training events specifically designed to upskill researchers in FAIR and Open Science. The organisation of the training series enabled a reflection on current challenges trainers face in making FAIR-by-design training resources and how to overcome them. The TRIPLE Training Toolkit is an open workflow for trainers to reproduce and adapt to organise training events following a FAIR-by-design method. It was created following the delivery of the TRIPLE Open Science Training Series. The purpose of the TRIPLE Training Toolkit is to provide effective support to the research community in the uptake and application of Open Science and FAIR Data management practices within training activities and to address the frequent findability and reusability issues related to the management of digital training materials. The Toolkit shows how the digital training materials created within the project are in line with the FAIR principles and enables for the experiment to be reproduced. It includes 11 reference documents referred to as reproducible templates that trainers can use and adapt to their needs along with illustrations of the process to facilitate the uptake of the method. The following files are deposited in Zenodo to serve as a reference for those wishing to reproduce this experiment within their own institution or for their own training activities. The first document to read is the README.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.005 | 0.006 |
| Open science | 0.004 | 0.008 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.266 | 0.201 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".