Digital Libraries that Demonstrate High Levels of Mutual Complementarity in Collection-level Metadata Give a Richer Representation of their Content and Improve Subject Access for Users
Notice bibliographique
Résumé
A Review of: Zavalina, O. L. (2013). Complementarity in subject metadata in large-scale digital libraries: A comparative analysis. Cataloging & Classification Quarterly, 52(1), 77-89. http://dx.doi.org/10.1080/01639374.2013.848316 Abstract Objective – To determine how well digital library content is represented through free-text and subject headings. Specifically to examine whether a combination of free-text description data and controlled vocabulary is more comprehensive than free-text description data alone in describing digital collections. Design – Qualitative content analysis and complementarity comparison. Setting – Three large scale cultural heritage digital libraries: one in Europe and two in the United States of America. Methods – The researcher retrieved XML files of complete metadata records for two of the digital libraries, while the third library openly exposed its full metadata. The systematic samples obtained for all three libraries enabled qualitative content analysis to uncover how metadata values relate to each other at the collection level. The researcher retrieved 99 collection-level metadata records in total for analysis. The breakdown was 39, 33, and 27 records per digital library. When comparing metadata in the free-text Description metadata element with data in four controlled vocabulary elements, Subject, Geographic Coverage, Temporal Coverage and Object Type, the researcher observed three types of complementarity: one-way, two-way and multiple-complementarity. The author refers to complementarity as “describing a collection’s subject matter with mutually complementary data values in controlled vocabulary and free-text subject metadata elements” (Zavalina, 2013, p. 77). For example, within a Temporal Coverage metadata element the term “19th century” would complement a Description metadata element “1850–1899” in the same record. Main Results – The researcher found a high level of one-way complementarity in the metadata of all three digital libraries. This was mostly demonstrated by free-text data in the Description element complemented by data in the controlled vocabulary elements of Subject, Geographic Coverage, Temporal Coverage, and Object Type. Only one library demonstrated a significant proportion (19%) of redundancy between free-text and controlled vocabulary metadata. An example of redundancy found included a repetition of geographic information in both a Description and Geographic Coverage metadata elements. Conclusion – The author reports high levels of mutual complementarity in the three cultural heritage digital libraries studied. The findings demonstrate that collection-level metadata which includes both free-text and controlled vocabulary is more representative of the intellectual content of the collections and improves subject access for users. The author maintains that there is no standard for collection-level metadata descriptions, and that this research may contribute to best practice guidelines in this area. It is unclear whether the digital libraries studied had written policies in place on how to describe collections and if those policies were adhered to in practice. The author expresses a need for further research to be conducted on collection-level metadata in other domains, such as science and interdisciplinary digital libraries, and on other scales (e.g., regional or state collections) and geographic regions beyond Europe and the United States.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,019 | 0,042 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,018 | 0,027 |
| Études des sciences et des technologies | 0,004 | 0,010 |
| Communication savante | 0,016 | 0,017 |
| Science ouverte | 0,001 | 0,016 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,010 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».