Bibliographic record
Abstract
Many factors can overwhelm the image searcher who is trying to retrieve images. Consequently, retrieving images is still a problem for the majority of individuals, especially for images associated with text written in different languages, unknown to users. The purpose of this exploratory study is to investigate the factors affecting search behaviours of image users and to examine how individuals formulate their queries to retrieve museum object images indexed in different languages, using different search engines. This study also compared the search behaviours using different types of search engines to retrieve specific museum object images. It also highlighted how multilingual search functionalities are perceived and used when performing an image search task in a multilingual retrieval context. Thirty participants randomly divided into three independent groups assigned to one different search engine was used for this study. Each participant was asked to retrieve three mages using an all-purpose search engine and a specialized search engine. Once the retrieval of the images was completed, the participants filled out a questionnaire to gather comments on the retrieval tasks they performed and information on their search behaviours of Web images indexed in different languages. Multilingualism plays a strategic role in the quality and effectiveness of communication services offered on the Internet. Consequently, it is of significant importance to make information available to the largest audience possible and to overcome language barriers by providing tools suited to the real and current needs of image searchers. The main contribution of this pilot study is to enhance the knowledge and understanding of image searching behaviour, in order to provide a basis for the modelling of a new search interface that takes into account the needs and expectations of real users.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".