Browse, search and serendipity
Notice bibliographique
Résumé
Large digital document collections ideally provide multiple routes into data imagined for different users and different use-cases: thematic and hierarchical (drill-down) browsability for casual users, and precisely-targeted complex search functionality to answer granular queries and generate subcollections for specific research purposes. Responding to recent critical work on digital editions and periodical print surrogates (e.g. Mussell 2012, 2016; Gooding), and on the visual interface as a form of graphic knowledge (Drucker), this chapter will examine the challenges in building a big tent digital project that anticipates users’ needs. The Digital Victorian Periodical Poetry Project (DVPP) has a particular interest in responding to this challenge, which is complicated by the nature of its own collection. The project’s methodological principles are based on poetry’s place on the periodical page, from the inclusion of periodical poem page scans (facsimile browser, poem page rendering), to the indexing protocols (designed around how contemporary periodical readers would understand poems and their illustrations), to encoding a representative sample of poems based on decadal years from 1820 to 1900 (including material as well as poetic features). But our approach to the front end application (facsimile browser, poem page rendering, index of poems and personography, digital edition, advanced search pages) is based around offering the user multiple ways to search and find material that moves away from the poem’s embedded periodical print origins, and even the conceptual and functional principles of the codex, to allow for complex and serendipitous discovery. The challenge of this digital project is to relate the project’s indexing and encoding principles to users’ anticipated research, particularly given the relationship between the index (c.15,500 poems across 21 long Victorian periodicals), personography (c.4,000 records for poets, illustrators and translators), and the TEI XML- encoded poem sample (c.2,000 poems and c. 11,000 lines of poetry). This chapter examines relationships between the underlying metadata and text-encoding, as well as the affordances DVPP will eventually offer the end-user. We conclude by offering guidelines based on building search interfaces that are useful to researchers. Firstly, we address practical problems. Enlarging project features can make interfaces potentially confusing, and expanding interdependencies can also produce incompatible features. Workflow is crucial: user discoverability is contingent on encoding, and yet predicting search parameters is contingent on a good understanding of data that only emerges as the project advances. We suggest a workflow where metadata structures and labels can be trivially revised, with the search and browse interfaces automatically adapted to such changes. Secondly, we turn to the conceptual imagining of the anticipated user, by comparing DVPP with cognate digital editions and commercial indexes and digital surrogates (such as those owned by ProQuest), to ask how digital editions can guide users to engage critically and actively with multiple methods of browse, search, and serendipitous discovery, rather than approaching search functionality as simply a means to an end.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,003 | 0,016 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,005 | 0,005 |
| Études des sciences et des technologies | 0,004 | 0,010 |
| Communication savante | 0,016 | 0,036 |
| Science ouverte | 0,001 | 0,013 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,029 | 0,008 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».