Bibliographic record
Abstract
Introduction Information retrieval systems are purposeful devices developed to service multiple objectives from locating the current weather conditions to identifying critical evidence for use in complex decision making. Those systems evolved from extracting bibliographic data held in large libraries of references, to retrieving nuggets of information from full-text repositories. However, we tend to think of information retrieval systems simply as generic search systems that respond to a query with a set of results to meet some information need, rather than purposeful applications whose raison d’êtr e is to deliver task-specific information that leads to problem resolution. The early discussions about information retrieval systems implicitly and erroneously equated the concept of task with a user's information need, problem, question or request (e.g. Saracevic et al., 1988; Saracevic and Kantor, 1988a, 1988b; Tague-Sutcliffe, 1992); the exact word used depended on the decade and origin of the work. The task that triggered the need was usually considered peripheral to the research. The explicit study of task is relatively new in information science. It appeared, for example, as an index term only in the second edition of Case's (2002, 2007) seminal analyses of information needs and seeking research. Its systematic use in information seeking and retrieval research is within the 21st century, a focus that occurred for several reasons, and possibly as a consequence of all of them: • Studies of ‘online searching’ went from being assessments of search outcomes and processes (see, for example, Saracevic et al., 1988; Saracevic and Kantor, 1988a, 1988b), to testing experimental variables (e.g. Koenemann and Belkin, 1996). Research protocols also went from being quasi-experimental and observational, to being more formally designed as human-based experiments following standard research design practices. • With the emergence of human–computer interaction in the 1980s came user-centred design and user testing, which brought user intentions, requirements and tasks to the forefront of systems development, and with it the premise that systems support people and the work that they do – their tasks. Using scenarios and tasks to design systems and for system and usability testing became de rigueur (see, for example, Rosson and Carroll, 2002). These developments have influenced information seeking and retrieval research, although a set of requirements has rarely been specified for a particular information retrieval application, for example a law office or an educational setting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.008 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.010 | 0.010 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.003 | 0.001 |
| Insufficient payload (model declined to judge) | 0.027 | 0.019 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".