Metadata and crowdsourced data for access and interaction in digitallibrary user interfaces
Bibliographic record
Abstract
Introduction Metadata has remained a major area of research in information science for nearly two decades. The increasing number and variety of metadata formats and standards has given rise to a number of digital library projects and initiatives that have focused on semantic interoperability among various metadata formats and standards. Interoperability, like metadata, is a widely researched and discussed topic in the literature of digital libraries. However, this chapter does not discuss interoperability per se; rather, it focuses on the use of metadata in the search interfaces of digital libraries. With the widespread use of metadata in digital libraries as access and retrieval points, it seems logical that they be used in user interfaces to support information seeking strategies. Shiri (2008) reported a study of metadata-enhanced visual interfaces and found that visual interfaces enhanced with metadata are an emerging category of visual interfaces. The growing number of digital libraries that create, maintain and support a variety of metadata provide ample opportunity for designers and developers of user interfaces. This chapter evaluates and compares four digital library user interfaces from four different countries (Edmonton Public Library (EPL), Canada; Trove: National Library of Australia, Australia; the Ann Arbor District Library System, USA; and the British Library, UK) in order to identify new developments in the use of metadata and to explore the emerging trends and new features and functionalities, such as social tags, recommendations, reviews and ratings in digital library user interfaces. The next section of the chapter introduces the definition, types and standards of metadata, followed by an introduction to digital libraries and user interfaces. Then the chapter presents the methodology used in the evaluation of the user interfaces of these four digital libraries, the findings and related discussion. Finally, the concluding section highlights some key trends and makes suggestions for future research. Metadata: definition, types and standards Numerous definitions of the term ‘metadata’ have been proposed by various research and development communities, including library and information science, archives, museums, computing, information technology, government organizations and educational institutions. This trend in itself points to the importance, popularity, usefulness and utility of metadata in various contexts, domains and disciplines. They all share the same philosophy that metadata aims to bring order to digital information and to support consistent and coherent description and discovery of digital objects.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.006 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.008 | 0.011 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.021 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".