What Makes Retrieval Optimal for Context-Level RAG? A Joint Top-k Approach
Abstract
Large language models can answer user question effectively, but they still suffer from hallucinations and outdated knowledge. Retrieval-Augmented Generation (RAG) has emerged as a popular solution by retrieving accurate and up-to-date reference documents from external databases and incorporating them into the generation context. Most existing theoretical formulations of RAG focus on the document-level setting, where candidate answers are generated for each retrieved document and then ranked. In practice, however, the more commonly deployed paradigm is context-level RAG, which directly generates a final answer conditioned on the top- retrieved documents. This creates a notable gap between existing theory and practical systems. To bridge this gap, we analyze context-level RAG through a probabilistic marginalization framework and derive a joint top- retrieval objective whose ranking criterion depends on the question, answer, and comprehensive context. We show that this joint top- objective is optimal for downstream answer generation. However, directly retrieving according to this criterion is infeasible in practice, since the answer and comprehensive context are unknown at retrieval time. To address this challenge, we propose JBR-C (Joint-Based Retriever for Context-level RAG), a practical retriever that uses listwise learning-to-rank to approximate the latent joint ranking signal with respect to the best answer and context. Extensive experiments on four question-answering benchmarks show that JBR-C consistently outperforms representative retrievers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.