RIT-RAG: Navigating Document Corpora with Retrieval-Induced Trees
Abstract
Retrieval-augmented generation (RAG) grounds language models in external corpora. Agentic RAG enables iterative search, yet exposes the model to isolated chunks without document structure, making it difficult to distinguish relevant evidence from chunks that merely resemble the query. Structure-aware methods such as PageIndex navigates using the structure but cannot scale to structures of large corpora that cannot fit in the LLM context. Hence they first commit to a single document using a document retriever and cannot recover from a wrong choice. We propose RIT-RAG (Retrieval-Induced Tree RAG), which combines content retrieval with structural navigation. Offline, RIT-RAG builds a tree for each document from its table of contents or sitemap. At query time, it retrieves a broad set of chunks and uses their positions to induce manageable sub-trees, potentially across multiple documents. An LLM agent navigates these sub-trees, selectively reads promising nodes, and reformulates queries when needed. Thus, retrieval proposes where to look, while the agent decides what to read. Across financial, scientific, and customer-support benchmarks, RIT-RAG achieves the highest answer accuracy among vanilla, graph-based, and agentic baselines. On EntDocs Benchmark, our new benchmark of 2.84 million technical-documentation webpages, it improves accuracy by 6.8–11.4 points over the strongest baseline across three LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.