DiSSCo Prepare Milestone report MS8.6 - Identifying Indicators for Alignment
Bibliographic record
Abstract
The Distributed System of Scientific Collections (DiSSCo) is part of an international landscape of bio-, geodiversity, environmental and life sciences related research infrastructures and organisations. This milestone seeks to provide context on DiSSCo’s current positioning within this landscape and to identify opportunities for future alignment and co-operation. The outputs from this report will help to inform future stakeholder engagement plans and prioritisation. A stakeholder analysis workshop identified the key stakeholders within this domain, with the Global Biodiversity Information Facility (GBIF), Catalogue of Life (CoL) Geoscience Collections Access Service (GeoCASe), Biodiversity Information Standards (TDWG) and the International Barcode of Life all classified as having high influence and interest in DiSSCo. The stakeholder analysis was supported by insight from the infrastructure contact zones survey, which is an analytic framework used to identify the possible synergies, complementarities and collaboration areas of DiSSCo with other organisations. This provides a systematic and standardised methodology for ranking a wide range of the activities of the different infrastructures relevant to the biodiversity informatics domain. LifeWatch and GBIF both aim to be more generalist infrastructures within the biodiversity informatics landscape, and operate across a large number of activities. There is contact between the activities of these organisations and DiSSCo, and enhanced collaborations may help to maximise the value of projects within this space. DiSSCo will also benefit from the expertise of more specialist infrastructures, such as GeoCASe and the Biodiversity Heritage Library.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.058 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.010 | 0.006 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.004 | 0.004 |
| Insufficient payload (model declined to judge) | 0.109 | 0.091 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".