Love notes to our future selves: Digital preservation and data curation
Notice bibliographique
Résumé
As the volume of research data grows in both size and complexity, concerns about maintaining access to files that are difficult (or impossible) to migrate, reliant on software that is not openly available or no longer accessible, or of such poor quality that they cannot be reused are heightened by an awareness of the environmental cost of digital storage. While archivists have a long history of practice guiding preservation decision-making and the deaccessioning or removal of records from archives, it is not clear how widely archival appraisal theory informs the approach to research data (Dorey, Hurley, and Knazook 2022). Long-term preservation of digital research data will prove challenging for repositories and preservationists, requiring substantially more information than is typically collected to support decisions about what to keep, how to maintain accessibility, and for how long. Preservationists need information about when the files were created, by whom, and using what software or tools, along with an understanding of the relevance of the data to the community of practice and its perceived long-term value. Curators play a critical role in communicating the informational value of datasets, and through their work with depositors, are in a unique position to collect information that will inform preservation decisions and reduce duplication of effort as data are (re)appraised over time, but they are often disconnected from preservation decision-making. In this panel discussion, we explore how training data curators in archival appraisal can help ensure long-term access to research data. We will hear an overview of a combined data curation and preservation workflow, and 3 institutions will discuss their experiences testing and refining a checklist developed at the Digital Research Alliance of Canada to record appraisal information about incoming datasets. We will end with a panel discussion and share a public version of the checklist.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,032 | 0,059 |
| Science ouverte | 0,004 | 0,009 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,005 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».