acceptodds
Under review as a conference paper at ICLR 2027

Generating Leakage-Free Benchmarks for Robust RAG Evaluation

Abstract

Retrieval-augmented generation (RAG) is widely used to augment large language models (LLMs) with external knowledge. However, many benchmark datasets, designed to test RAG performance, comprise many questions that can already be answered from an LLM’s parametric memory. This leads to unreliable evaluation. We refer to this phenomenon as knowledge leakage—cases where RAG tasks are solvable without retrieval. This issue worsens over time due to benchmark aging. As benchmarks are reused for training, their contents are increasingly absorbed into model parameters, making them less effective for evaluating retrieval. We introduce SeedRG, a semi-synthetic benchmark generation pipeline that mitigates knowledge leakage and addresses the issue of benchmark aging. Starting from a seed benchmark dataset, SeedRG extracts a reasoning graph from question–context pairs to capture their underlying reasoning structure, and then generates new examples via type-constrained entity replacement. This process produces structurally similar but novel instances that are unlikely to exist in the model’s parametric knowledge, while preserving the original reasoning patterns. To ensure quality, we incorporate two verification steps: (1) a reasoning-graph consistency check to maintain task difficulty, and (2) a knowledge-leakage filter to exclude instances answerable without retrieval. We evaluate SeedRG on three seed benchmarks (HotpotQA, 2WikiMultihopQA, QASC) and three popular LLMs (GPT-5, Claude Sonnet 4.5, Gemini 2.5 Flash). SeedRG reduces knowledge leakage by at least 78% while preserving reasoning difficulty. By removing the confounding effect of parametric knowledge, SeedRG reveals meaningful variability across RAG systems that is otherwise obscured. Prior benchmarks show uniformly high performance across RAG methods (HippoRAG, GraphRAG, OGRAG, SemanticRAG), because performance is dominated by model knowledge. In contrast, SeedRG surfaces clear differences in retrieval and reasoning ability across the methods. Beyond benchmark construction, we provide a systematic analysis linking reasoning difficulty to graph structure, showing how structural variations induce predictable changes in model accuracy. Together, these results demonstrate that SeedRG enables more discriminative and robust evaluation of RAG systems.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.