MétaCan
Menu
Retour à la cohorte
Enregistrement W4414602742 · doi:10.28995/2073-0101-2025-3-745-760

Digital archives: problems of definition, systematization and boundaries of the research field. The end of the 20th – the first quarter of the 21st century

2025· article· en· W4414602742 sur OpenAlexaboutno aff
Natalia Dushakova, M. D. Shkil

Notice bibliographique

RevueHerald of an Archivist · 2025
Typearticle
Langueen
DomaineComputer Science
ThématiqueResearch Data Management Practices
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésProblematizationQuarter (Canadian coin)Digital preservationArchival scienceTerm (time)Digital RevolutionCultural heritageDigitization

Résumé

récupéré en direct d'OpenAlex

The article is devoted to the problematization and clarification of the term "digital archive", the analysis of priority areas in the research of digital archives and the issues of systematization of disparate sources in the digital space. As a result of the digital revolution and the general growth of interest in the preservation of historical heritage, new ways of processing and storing information have become widespread. Not only archives and libraries, but also scientific and educational organizations, specialists from various fields, as well as ordinary people have joined the creation and popularization of digital archives. In this context, archivists and historians were not so much interested in the definition of "digital archive", which is important for theoretical understanding, as in the specifics of working with electronic documents, ensuring their safety and use, the problems of finding approaches to the study of initially digital sources, the preservation and popularization of archival heritage in a digital environment. At the same time, many researchers (besides archivists and historians, anthropologists, folklorists, geographers, cultural scientists, historians of science, etc.) have started developing databases on topics that are close to them, which even colleagues from the same or related disciplines do not always know about. The current definitions of the term "digital archive" seem too general and vague. As a result, digital archives generally include all stored digital objects of some significance, any electronic documents, and even in the broadest sense, a social network or the entire Internet. On the other hand, the general, at first glance, definition of a digital archive as digitized collections of documents ignores initially digital documents. Taking into account the accumulated research experience, based on the analysis of the existing historiography on the problem and the identified gaps in understanding digital archives, the authors propose in the article to clarify the definition of the concept of "digital archive", drawing attention to several important circumstances: first, the need to cover both digitized and initially digital documents with this concept.; Secondly, the fact that digital documents undergo a selection procedure before becoming part of a digital archive (there will be no raw funds or random documents); thirdly, the feature of a digital archive is the presentation of documents in a systematic form. Of course, in creating a digital archive, an important role is played by the point of view when selecting documents and how they are systematized, and this is always subjective, which can both impose restrictions on their use and open up additional opportunities. The authors formulate the definition of a "digital archive" as a set of digitized or digitally created documents selected for storage and presented in a systematic way. One of the possible solutions to the problem of the fragmentation of digital sources (both digitized and initially digital) is the creation and publicly available publication of a consolidated catalog of digital archives, which is being developed by a team of authors at the Archive of the Russian Academy of Sciences.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,035
score de la tête « metaresearch » (Gemma)0,059
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCommunication savante
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Synthèse · Signal consensuel: Synthèse
Score de désaccord entre enseignants0,960
Score d'incertitude au seuil0,187

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0350,059
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0080,013
Études des sciences et des technologies0,0130,073
Communication savante0,0400,040
Science ouverte0,0030,018
Intégrité de la recherche0,0070,011
Charge utile insuffisante (le modèle a refusé de juger)0,0070,002

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,038
Tête enseignante GPT0,297
Écart entre enseignants0,259 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueHerald of an ArchivistMême sujetResearch Data Management PracticesTravaux en français237 207