Towards Robust Retrieval-Augmented Generation Based on Knowledge Graph: A Comparative Analysis
Bibliographic record
Abstract
Retrieval-Augmented Generation (RAG) was first introduced to enhance the capabilities of Large Language Models (LLMs) beyond their encoded-prior knowledge. This is achieved by providing LLMs with an external source of knowledge, which helps to reduce factual hallucinations and enables the access to new information, typically not available during their pretraining phase. Despite its benefits, there is an increasing concern with the impact of inconsistent retrieved information towards LLMs’ responses. Hence, the Retrieval-Augmented Generation Benchmark (RGB) was introduced as a new testbed for RAG evaluation, meant to assess the robustness of LLMs towards inconsistency in the retrieved information. In this work, we use the RGB corpus to evaluate LLMs in four scenarios: (1) noise robustness; (2) information integration; (3) negative rejection; and (4) counterfactual robustness. We perform a comparative analysis between the RAG baseline defined by the RGB and variations of GraphRAG, which is a RAG system based on a Knowledge Graph (KG) and developed to retrieve relevant information from large documents. We tested GraphRAG under three customization to improve its robustness. Our approach demonstrates improvements compared to the RGB baseline, providing insights on how to design more reliable RAG systems, tailored for real-world scenarios.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.010 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".