Digital library reference model - in a nutshell
Bibliographic record
Abstract
The Digital Library universe is a complex framework bringing together many disciplines and fields, spanning data management, information retrieval, library sciences, document management, information systems, web image processing, artificial intelligence, human-computer interaction and digital curation. The Digital Library universe is also an interplay of professional roles, encompassing cataloguing and curating, defining, customising and maintaining the Digital Library and its services, as well as developing and customising software. Such complexity and diversity in terms of approaches, solutions and systems has driven the need for common foundations that foster best practices and help focus further advancement in the field. The Digital Library Reference Model aims at contributing to the creation of such foundations. It is the result of a collective understanding on Digital Libraries that has been acquired by European research groups under the umbrella of the European funded DELOS Network of Excellence on Digital Libraries, as well as the international scientific community active in the field of Digital Libraries. The outcomes of DELOS have been taken forward by DL.org, a project funded by the Cultural Heritage and Technology Advanced Learning Unit of the Information Society Directorate-General of the European Commission, working in synergy with a team of international experts in the field to enhance and extend the Reference Model. The Digital Library Reference Model is a conceptual framework aimed at capturing significant entities and their relationships in the digital library universe with the goal of developing a concrete model of it. This conceptual framework can be exploited for coordinating approaches, solutions and systems development in the digital library area. In particular, it is envisaged that in the future Digital Library ‘systems’ will be described, classified and measured according to the key elements introduced by this model. Behind the Reference Model’s efforts there has been the driving force of The Digital Library Manifesto, laying down the main notions characterising the Digital Library universe in rather abstract terms.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.007 | 0.009 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.011 | 0.021 |
| Open science | 0.005 | 0.006 |
| Research integrity | 0.004 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.015 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".