Bridging the gap in digital art preservation: interdisciplinary reflections on authenticity, longevity and potential collaborations
Bibliographic record
Abstract
Digital casualties: challenges for digital art preservation Born-digital art is fundamentally art produced and mediated by a computer. It is an art form within the more general ‘media art’ category (Altshuler, 2005; Paul, 2008a, 2008b; Depocas et al., 2003; Grau, 2007; Graham and Cook, 2010; Lieser, 2010) and includes software art, computer-mediated installations, internet art and other heterogeneous art types. The boundaries of digital art are particularly fluid, as it merges art, science and technology to a great extent. The technological landscape in which digital art is created and used challenges its long-term accessibility, the potentiality of its integrity, and the likelihood that it will retain authenticity over time. Digital objects – including digital art works – are fragile and susceptible to technological change. We must act to keep digital art alive, but there are practical problems associated with its preservation, documentation, access, function, context and meaning. Preservation risks for digital art are real: they are technological but also social, organizational and cultural. Digital and media art works have challenged ‘traditional museological approaches to documentation and preservation because of their ephemeral, documentary, technical, and multi-part nature’ (Rinehart, 2007b, 181). The technological environment in which digital art lives is constantly changing, which makes it very difficult to preserve this kind of art work. All art is subject to change. This can occur at art object level and at context level. In most circumstances change is very slow, but in digital art this isn't the case anymore because it is happening so quickly, owing to the pace of technological development. Surely the increased pace of technological development has more implications than just things happening faster. Digital art, in particular, questions many of the most fundamental assumptions of the art world: what is a work of art in the digital age? What should be retained for the future? Which aspects of a given work can be changed and which must remain fixed for the work to retain the artist's intent? How do museums collect and preserve? Is a digital work as fragile as its weakest components? What is ownership? What is the context of digital art? What is a viewer? It is not feasible for the arts community to preserve over the centuries working original equipment and software.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.028 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.020 | 0.061 |
| Scholarly communication | 0.030 | 0.046 |
| Open science | 0.003 | 0.029 |
| Research integrity | 0.006 | 0.008 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".