MétaCan
Menu
Back to cohort
Record W7108325384 · doi:10.5281/zenodo.17725564

Write it Down! Fostering Responsible Reuse of Cultural Heritage Data with Interoperable Dataset Descriptions

2025· article· enc· W7108325384 on OpenAlexaboutno aff

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2025
Typearticle
Languageenc
FieldArts and Humanities
TopicDigital Humanities and Scholarship
Canadian institutionsnot available
Fundersnot available
KeywordsCultural heritageInteroperabilityDocumentationTransparency (behavior)ReuseContext (archaeology)Stewardship (theology)Linked data

Abstract

fetched live from OpenAlex

Abstract Cultural heritage institutions have seen a surge in the creation of datasets ready for computational use, while researchers increasingly experiment with datasets through computational processing and AI-assisted methods. For both groups, issues of transparency have sparked interest in developing documentation practices cutting across the artificial intelligence/machine learning (AI/ML) and the digital cultural heritage (DCH) sector, aiming to provide better information on e.g., the purpose, composition, reusability, collection processes and provenance, or societal biases reflected in datasets. The publication of Datasheet for Datasets (Gebru et al., 2021) and the Collections as Data movement (Padilla et al. 2023) have sparked the definition of guidelines for dataset creators and publishers who want to follow FAIR and CARE principles and make it easier for one to reuse their data in a responsible, well-informed manner. Gathering CH professionals, technical experts and humanities scholars from the Europeana Research and EuropeanaTech communities, the Datasheets for Digital Cultural Heritage working group has adapted existing ML documentation approaches to the DCH case. As a first outcome, a template (Alkemade et al, 2023) has sought to address the complexities of DCH datasets, shaped by layered curatorial decisions, often subject to evolving and non-linear trajectories. In the spirit of the common European data space for cultural heritage (2025), which is being deployed under the stewardship of the Europeana Initiative, the working group has then supported professionals interested in applying the template in their institutional context (see for example Lehmann et al., 2024) and fostered exchanges with other initiatives emerging at the European level and exploring suitable ways to describe datasets. One key initiative in this regard concerns the proposal for Data-Envelopes for Cultural Heritage (Luthra et al., 2024), which has focused specifically on providing machine-readable descriptions of datasets, especially considering the W3C Data Catalogue Vocabulary (DCAT) that is used in many data portals. The goal of this collaboration is both to validate and further refine the existing templates following a community-led approach, and to investigate how to ensure (human-machine) interoperability in the data space, which aims to establish a diverse data offer (including datasets suitable for AI applications, as illustrated by the AI4Culture platform (2025)) as well as making use of DCAT. Our contribution will report on the following ongoing work: Alignment with DCAT: DCH datasheet fields are being mapped to DCAT to enable machine-readability Alignment between DCH datasheets and data-envelopes, establishing conceptual and structural compatibility, and supporting future integration with other legal, technical and ethical frameworks. Gathering a set of exemplary dataset descriptions Creation of (prototype) tooling to support and simplify the creation, reuse and integration of descriptions into existing workflows. We also plan to discuss new items that will begin before the conference: Identify possible connections with data research plans and data management plans. This may extend to interoperability with emerging European Cultural Heritage Cloud (ECHOES, 2025). Establish a modular structure for descriptions, aiming at operationalising the templates by defining building blocks, including a ‘core’ common to most DCH collections and a series of ‘profiles’, tailored to research data management and AI/ML workflows (e.g., AI Model Research Documentation Sheet (AIRDocS) (Oberbichler, 2025) Providing guidance to use these modules and possibly develop custom ones. While some components remain under active development (e.g. prototype, profiles and guidelines for their development), we present this work in progress to foster dialogue and invite broader engagement from the Fantastic Futures community. References AI4Culture project (2025). AI4Culture, Empowering Cultural Heritage through Artificial Intelligence. https://ai4culture.eu Alkemade, H., Claeyssens, S., Colavizza, G., Freire, N., Irollo, A., Lehmann, J., Neudecker, C., Osti, G., & van Strien, D. (2023, September 25). Datasheets for Digital Cultural Heritage Datasets—Template v.1. Zenodo. https://zenodo.org/records/8375034 Alkemade, H., Claeyssens, S., Colavizza, G., Freire, N., Lehmann, J., Neudecker, C., Osti, G., & Van Strien, D. (2023). Datasheets for Digital Cultural Heritage Datasets. Journal of Open Humanities Data, 9, 17. https://doi.org/10.5334/johd.124 Common European data space for cultural heritage (2025), Welcome to the Common European data space for cultural heritage. https://www.dataspace-culturalheritage.eu/en ECHOES project (2025), ECCCH, The Cultural Heritage Cloud, https://www.echoes-eccch.eu/ Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for Datasets. Communications of the ACM, 64(12), 86–92. https://doi.org/10.1145/3458723 Luthra, M., & Eskevich, M. (2024). Data-Envelopes for Cultural Heritage: Going beyond Datasheets. In I. Siegert & K. Choukri (Eds.), Proceedings of the Workshop on Legal and Ethical Issues in Human Language Technologies @ LREC-COLING 2024 (pp. 52–65). ELRA and ICCL. https://aclanthology.org/2024.legal-1.9 Lehmann, J., & Schneider, S. (2024). Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB). https://doi.org/10.5281/zenodo.13284442 Oberbichler, S. (2025). AI Model Research Documentation Sheet (AIRDocS). https://doi.org/10.5281/zenodo.15046713 Padilla, T., Scates Kettler, H., Varner, S., & Shorish, Y. (2023). Vancouver Statement on Collections as Data. https://zenodo.org/records/8342171 Pushkarna, M., Zaldivar, A., & Kjartansson, O. (2022). Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI. 2022 ACM Conference on Fairness, Accountability, and Transparency, 1776–1826. https://doi.org/10.1145/3531146.3533231 World Wide Web Consortium. (2024). Data Catalog Vocabulary (DCAT) - Version 3. https://www.w3.org/TR/vocab-dcat-3/

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.001
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScience and technology studies, Scholarly communication, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.810
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.001
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0030.001
Scholarly communication0.0060.002
Open science0.0030.008
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0220.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.167
GPT teacher head0.294
Teacher spread0.128 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicDigital Humanities and ScholarshipFrench-language works237,207