MétaCan
Menu
Retour à la cohorte
Enregistrement W4409919791 · doi:10.62637/sup.ghst9020.6

Browse, search and serendipity

2025· book-chapter· en· W4409919791 sur OpenAlexaff
Alison Chapman, Martin Holmes, Kaitlyn Fralick, Kailey Fukushima, Narges Montakhabi Bakhtvar, Sonja Pinto

Notice bibliographique

Revuenon disponible
Typebook-chapter
Langueen
DomaineArts and Humanities
ThématiqueDigital Humanities and Scholarship
Établissements canadiensUniversity Canada WestUniversity of Victoria
Organismes subventionnairesArts and Humanities Research Board
Mots-clésPoetryComputer scienceSerendipityCasualWorld Wide WebRendering (computer graphics)Table of contentsFacsimileIndex (typography)Information retrievalComputer graphics (images)ArtLiterature

Résumé

récupéré en direct d'OpenAlex

Large digital document collections ideally provide multiple routes into data imagined for different users and different use-cases: thematic and hierarchical (drill-down) browsability for casual users, and precisely-targeted complex search functionality to answer granular queries and generate subcollections for specific research purposes. Responding to recent critical work on digital editions and periodical print surrogates (e.g. Mussell 2012, 2016; Gooding), and on the visual interface as a form of graphic knowledge (Drucker), this chapter will examine the challenges in building a big tent digital project that anticipates users’ needs. The Digital Victorian Periodical Poetry Project (DVPP) has a particular interest in responding to this challenge, which is complicated by the nature of its own collection. The project’s methodological principles are based on poetry’s place on the periodical page, from the inclusion of periodical poem page scans (facsimile browser, poem page rendering), to the indexing protocols (designed around how contemporary periodical readers would understand poems and their illustrations), to encoding a representative sample of poems based on decadal years from 1820 to 1900 (including material as well as poetic features). But our approach to the front end application (facsimile browser, poem page rendering, index of poems and personography, digital edition, advanced search pages) is based around offering the user multiple ways to search and find material that moves away from the poem’s embedded periodical print origins, and even the conceptual and functional principles of the codex, to allow for complex and serendipitous discovery. The challenge of this digital project is to relate the project’s indexing and encoding principles to users’ anticipated research, particularly given the relationship between the index (c.15,500 poems across 21 long Victorian periodicals), personography (c.4,000 records for poets, illustrators and translators), and the TEI XML- encoded poem sample (c.2,000 poems and c. 11,000 lines of poetry). This chapter examines relationships between the underlying metadata and text-encoding, as well as the affordances DVPP will eventually offer the end-user. We conclude by offering guidelines based on building search interfaces that are useful to researchers. Firstly, we address practical problems. Enlarging project features can make interfaces potentially confusing, and expanding interdependencies can also produce incompatible features. Workflow is crucial: user discoverability is contingent on encoding, and yet predicting search parameters is contingent on a good understanding of data that only emerges as the project advances. We suggest a workflow where metadata structures and labels can be trivially revised, with the search and browse interfaces automatically adapted to such changes. Secondly, we turn to the conceptual imagining of the anticipated user, by comparing DVPP with cognate digital editions and commercial indexes and digital surrogates (such as those owned by ProQuest), to ask how digital editions can guide users to engage critically and actively with multiple methods of browse, search, and serendipitous discovery, rather than approaching search functionality as simply a means to an end.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,003
score de la tête « metaresearch » (Gemma)0,016
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesCommunication savante
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Théorique ou conceptuel · Signal consensuel: Théorique ou conceptuel
GenreSignal candidat: Synthèse · Signal consensuel: aucune
Score de désaccord entre enseignants0,984
Score d'incertitude au seuil0,097

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0030,016
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0050,005
Études des sciences et des technologies0,0040,010
Communication savante0,0160,036
Science ouverte0,0010,013
Intégrité de la recherche0,0020,002
Charge utile insuffisante (le modèle a refusé de juger)0,0290,008

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,076
Tête enseignante GPT0,234
Écart entre enseignants0,157 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeThéorique ou conceptuel
Domainenon disponible
GenreSynthèse

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même sujetDigital Humanities and ScholarshipTravaux en français237 207