Digital information interaction as semantic navigation
Bibliographic record
Abstract
Introduction In this chapter we focus on the research area of digital information interaction, which emphasizes searchers’ direct engagement with and manipulation of information objects as they search and browse through digital information environments. This is an area of active research that has opened up in recent years as information retrieval (IR) research has expanded its focus from the mechanics of retrieval (i.e. indexing, data structures and retrieval algorithms) to include a broader ‘retrieval in context’ perspective that takes into account the whole system, the affective, cognitive and physical attributes of users and the environment in which searching takes place (Ingwersen and Jarvelin, 2005). A number of meetings and workshops have focused on this area, including the Information Retrieval in Context (IRiX) workshops at the ACM SIGIR (Association for Computing Machinery Special Interest Group Information Retrieval) conference (2004–5), the Information Interaction in Context (IIiX) Conference (2006–ongoing) and the Human Computer Information Retrieval (HCIR) Workshops (2007–ongoing). Immersive information systems, in which users are placed in content-rich environments with tools and technologies designed to help them search, filter and navigate through the information landscape, are now commonplace. In these types of environments, IR can no longer best be modelled as a transaction, but is more akin to an ‘experience’, in which searchers are influenced by the information objects they encounter and also shape and create their information environment by selecting, linking, tagging and commenting (Marchionini, 2008; Toms, 2002). For many people, these kinds of interactions with digital information are the primary means by which they read and learn, whether at work, at school or at leisure. However, numerous studies that utilize server logs of bibliographic and web based information systems have shown that digital information behaviour is markedly different from that in traditional print information environments. Digital information use tends to be non-linear, shallow and piecemeal. Readers jump frequently from one document to the next, ‘skimming and bopping’ through networked information spaces (JISC, 2008; Nicholas et al., 2007). Concerns have been raised that this type of information behaviour represents the ‘intellectual equivalent of empty calories’ (Rich, 2008) and inhibits deep, critical engagement with informational content (Marshall, 2005).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.004 |
| Science and technology studies | 0.001 | 0.005 |
| Scholarly communication | 0.007 | 0.011 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.019 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".