Do Not Blame Your RAG System: The Multi-Hop QA Benchmarks Need Revision
Abstract
Retrieval-augmented generation (RAG) in modern agents underpins knowledge, skill, memory and tool retrieval. RAG systems are primarily evaluated on three multi-hop question answering (QA) benchmarks, namely HotpotQA, 2WikiMultiHopQA and MuSiQue, which test whether a system can retrieve the relevant documents and generate the correct answer by reasoning over them. However, we show that these benchmarks have significant flaws. On average, 39.8% of the samples carry at least one defect; for example, the question cannot be answered as written, the gold documents do not support the official answer, or the official answer is wrong. Our preliminary studies further show that even when given the gold documents, models of different sizes and families still miss the official answer on 15.9% to 24.8% of all questions, mainly on the defective samples. These defects inherently bias the evaluation of RAG systems, which rests on a flawed standard that requires urgent revision. We target this goal and release two audited versions of each benchmark: a repaired version (22,255 samples in total) based on our LLM auditing framework and a human-verified subset (2,732 samples in total). Notably, by auditing all 22,398 samples in the development sets of the three benchmarks, our LLM Auditor flags 35.6% of HotpotQA, 26.8% of 2WikiMultiHopQA and 57.1% of MuSiQue as defective under a taxonomy of twelve defect labels we define. We further evaluate the state-of-the-art RAG methods on our audited benchmarks to set up a reliable baseline for the community. We hope these released benchmarks and our framework will enable more reliable evaluation and development of future RAG systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.