Improving the Performance of Arabic Information Retrieval Systems: The Issue of Resolving Word Sense Disambiguation
Bibliographic record
Abstract
This study aimed at assessing the performance and efficacy of the retrieval information (IR) systems implemented in three widely used search engines (Google, Bing, and Yahoo), specifically with regard to the challenge of word sense disambiguation in Arabic texts. Such a challenge has been confirmed to negatively influence the retrieval of the most relevant documents. Therefore, we extended the paradigm of using computational methods and natural language processing (NLP) tools, primarily tailored for processing English texts, to explore morphosyntactic as well as lexical issues disturbing the accuracy of Arabic IR systems. Findings revealed striking disparities in the efficacy of IR systems integrated into these search engines, which can be attributed to four principal challenges: (a) the intricate morpho-syntactic structures inherent in Arabic; (b) the idiosyncratic orthographical system of the Arabic script; (c) the multifaceted semantic flexibility of certain lexical elements; and (d) the intriguing diaglossic nature of Arabic, allowing for the coexistence of multiple linguistic varieties within a single discourse situation. Drawing from these findings, a series of solutions rooted in supervised machine learning techniques, including clustering models and adaptations based on geographic locations, are proposed. Moreover, the study advocates for the capacity of search engines to interpret queries across all Arabic varieties, encompassing vernacular dialects. Furthermore, the importance of search engines accommodating queries irrespective of the specific language adopted by users is underscored. While the research primarily centers on Arabic, its implications resonate beyond this language alone. By applying computational methodologies originally designed for English to Arabic, the study not only addresses the challenges specific to Arabic IR systems but also contributes valuable insights that transcend linguistic boundaries. Through a comparative lens, issues like word sense disambiguation between Arabic and English are juxtaposed, extracting lessons that can inform advancements in information retrieval for both languages.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".