Bibliographic record
Abstract
Introduction: a European approach? After considering the title of this chapter for some time a question presented itself. Is there such a thing as ‘a European approach’? How could this possibly be the case? Digital materials are consistent throughout the world. At least, they are consistent in their diversity: from New Zealand to Canada, to Brazil and Japan, these countries and their public and private organizations are creating a great mass of materials whose form can be recognized the world over. So can there be a ‘European approach’? Yes. Despite the electronic medium passing through boundaries across the world globalizing practices, there are certain attitudes, positions and circumstances that have outlined what can indeed be termed a ‘European approach’. Of course, the examples given below do not necessarily reflect a uniform European approach, but there is more than a hint of a ‘knowledge-space’ that exists in Europe contained in them. This knowledge-space is the product of political boundaries and histories, knowledge, awareness and terminology, and is one of the major contributing factors to the shape of digital preservation work in Europe. Additionally, there is international recognition that the progress that Europe has made provides elements of an approach for others. Through the exploration of the European political climate and a discussion on some major shifts in terminology a uniqueness should appear, an indefinable ‘something’ that has come into being to allow us to say that there is a European approach to digital preservation which is distinctly recognizable. For those not convinced of this view, there is a second angle to the chapter title. The title could also read: ‘Some other European approaches’. What is apparent throughout all the literature on issues surrounding digital preservation is that there is a recycling of examples. Vitality is required to keep ideas fresh and momentum going, and the approaches discussed here should offer new perspectives to some of the issues that have been discussed in many fora. Outside the cultural heritage space there are companies and institutions dealing with the persistence of their digital assets as a matter of organizational and business need. Their experiences are instructive to the debates taking place in conferences, journals, newsletters and newsgroups. If nothing else, the examples below will expand horizons within digital preservation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.005 | 0.008 |
| Scholarly communication | 0.009 | 0.008 |
| Open science | 0.001 | 0.005 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.011 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".