DictRAG: Composable Corpus Access for Multi-Hop Question Answering
Abstract
Multi-hop question answering requires combining constraints to locate evidence scattered across a corpus. Ranked retrieval returns limited passage lists, restricting subsequent filtering to retrieved candidates. We introduce DictRAG, which enables agents to compose corpus-level conditions before selecting evidence to read. A bounded SQL interface exposes documents as field–value records and supports exact lexical filtering, set operations, and reusable relation definitions across turns. An indexed backend executes these operations outside model context, returning only selected observations. In corpus-constrained evaluations on HotpotQA, 2WikiMultiHopQA, MuSiQue, and Bamboogle, DictRAG with Qwen3.5-27B achieves a dataset-macro average of 65.70% exact match (EM) and 75.98% F1. Average gains over both agentic retrievers also hold with Qwen3.5-9B and Gemma4-31B. Progressive ablations on MuSiQue show larger observed gains from exact search and condition composition than from cross-turn relation reuse. Under corpus expansion, DictRAG has higher observed EM and F1 than both retrievers on all three scaling datasets at 100K and 500K documents. With a expansion of the HotpotQA corpus to 5.23 million documents, DictRAG retains 65.50% EM and shows a smaller observed decline than both agentic retrievers: 3.0 percentage points, versus 5.2 for Agentic-BM25 and 4.1 for Agentic-BGE. These results support composable corpus access as a useful capability for multi-hop question answering over large document collections.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.