Utilizing Large Language Model for Conversational Information Seeking via Dual-Query Generation and Joint-Encoding
Bibliographic record
Abstract
Conversational retrieval leverages multi-turn conversations to meet users’ information needs, and accurately understanding the new intent has become a significant challenge in this field. Recently, the language comprehension and reasoning capabilities of large language models (LLMs) offer a viable solution to these challenges. In this article, we propose a new Dual-Query Generation and Joint-Encoding method by utilizing LLM for Conversational Information Seeking, abbreviated as DQ-CIS. Specifically, we propose a dual-query generation approach that leverages both open source and closed source LLMs to generate two complementary queries: a full-rewrite query that preserves the context semantics of the conversation and a condensed-rewrite query that emphasizes the core intent of the current query. Additionally, to better express the semantic information of the query, we propose a dual-query joint-encoding method, which enhances the thematic expression of query vectors by treating the dual-query as semantic complementary. A query coverage fine-tuned semantic matching method is also introduced to improve result relevance and ranking by fine-tuning the original retrieval scores by ColBERT. We conducted a number of experiments on seven publicly available conversational retrieval datasets. The results show that compared with other models, DQ-CIS has strong competitiveness in both retrieval efficiency and retrieval results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.006 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".