Inedxing as Problem Solving: A Cognitive Approach to Consistency
Bibliographic record
Abstract
Indexers differ in their judgment as to which terms reflect adequately the content of a document. Studies on interindexers' consistency identified several factors associated with low consistency, but failed to provide a comprehensive model of this phenomenon. Our research applies theories and methods from cognitive psychology to the study of indexing behavior. From a theoritical standpoint, indexing is considered as a problem solving situation. To access to the cognitive processes of indexers, three kinds of verbal reports are used. We will present results of an experiment in which four experienced indexers indexed the same documents. It will be shown that the three kinds of verbal reports provide complementary data on strategic behavior, and that it is of prime importance to consider the indexing task as an ill-defined problem, where the solution is partly defined by the indexer him(her)self.Résumé: Afin de mieux cerner l'origine des différences relevées par les études sur la cohérence inter-indexeurs, nous adjoignons les théories et méthodes issues de la psychologie cognitive (spécifiquement celles qui considèrent une tâche comme une situation de résolution de problème) à l'étude des processuscognitifs impliqués lors de l'indexation. Cette approche permet d'identifier les stratégies cognitives élaborées et utilisées par les indexeurs. Afin d'avoir accès à leurs processus cognitifs, trois types de verbalisations sont recueillis. Nous présenterons les résultats d'une expérimentation pour laquelle quatre indexeurs expérimentés ont analysé les mêmes documents. Les résultats avancés démontreront la complémentarité des données issues des trois types de verbalisations et l'importance de considérer l'indexation en tant que problème mal- défini; la solution étant définie en partie par l'indexeur.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.003 | 0.014 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".