Browse, search and serendipity
Bibliographic record
Abstract
Large digital document collections ideally provide multiple routes into data imagined for different users and different use-cases: thematic and hierarchical (drill-down) browsability for casual users, and precisely-targeted complex search functionality to answer granular queries and generate subcollections for specific research purposes. Responding to recent critical work on digital editions and periodical print surrogates (e.g. Mussell 2012, 2016; Gooding), and on the visual interface as a form of graphic knowledge (Drucker), this chapter will examine the challenges in building a big tent digital project that anticipates users’ needs. The Digital Victorian Periodical Poetry Project (DVPP) has a particular interest in responding to this challenge, which is complicated by the nature of its own collection. The project’s methodological principles are based on poetry’s place on the periodical page, from the inclusion of periodical poem page scans (facsimile browser, poem page rendering), to the indexing protocols (designed around how contemporary periodical readers would understand poems and their illustrations), to encoding a representative sample of poems based on decadal years from 1820 to 1900 (including material as well as poetic features). But our approach to the front end application (facsimile browser, poem page rendering, index of poems and personography, digital edition, advanced search pages) is based around offering the user multiple ways to search and find material that moves away from the poem’s embedded periodical print origins, and even the conceptual and functional principles of the codex, to allow for complex and serendipitous discovery. The challenge of this digital project is to relate the project’s indexing and encoding principles to users’ anticipated research, particularly given the relationship between the index (c.15,500 poems across 21 long Victorian periodicals), personography (c.4,000 records for poets, illustrators and translators), and the TEI XML- encoded poem sample (c.2,000 poems and c. 11,000 lines of poetry). This chapter examines relationships between the underlying metadata and text-encoding, as well as the affordances DVPP will eventually offer the end-user. We conclude by offering guidelines based on building search interfaces that are useful to researchers. Firstly, we address practical problems. Enlarging project features can make interfaces potentially confusing, and expanding interdependencies can also produce incompatible features. Workflow is crucial: user discoverability is contingent on encoding, and yet predicting search parameters is contingent on a good understanding of data that only emerges as the project advances. We suggest a workflow where metadata structures and labels can be trivially revised, with the search and browse interfaces automatically adapted to such changes. Secondly, we turn to the conceptual imagining of the anticipated user, by comparing DVPP with cognate digital editions and commercial indexes and digital surrogates (such as those owned by ProQuest), to ask how digital editions can guide users to engage critically and actively with multiple methods of browse, search, and serendipitous discovery, rather than approaching search functionality as simply a means to an end.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.016 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.004 | 0.010 |
| Scholarly communication | 0.016 | 0.036 |
| Open science | 0.001 | 0.013 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.029 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".