acceptodds
Under review as a conference paper at ICLR 2027

StageRAG: Answer-Supervised Cascade Optimization for Retrieval-Augmented Generation

Abstract

Retrieval-augmented generation (RAG) is commonly deployed as a cascade: a retriever recalls candidates, a reranker selects a compact context, and a language model generates the answer. Optimizing this cascade from answer feedback alone faces two challenges. Corpus-specific supporting chains are costly and ambiguous to annotate, creating a Gold-Evidence Bottleneck; meanwhile, final-answer feedback does not reveal where evidence was lost, creating Cascade Credit Ambiguity. We introduce StageRAG, an answer-supervised framework that improves both the retriever and reranker without gold supporting-document annotations. StageRAG records candidate progression in a logged RAG execution, uses an answer-and-citation verifier to determine answer correctness and support attribution, converts the resulting states into query-level preferences, and balances them with per-query caps. We also formalize a pipeline-aware ranking objective, derive a pointwise rule under idealized conditions, and bound its regret under bounded evidence interactions. On HotpotQA, 2WikiMultiHopQA, and TAT-QA, StageRAG improves F1 over the strongest compared answer-supervised baseline by 5.65, 12.44, and 6.43 points, respectively. Cascade diagnostics show improved alignment with downstream trajectory preferences and recovery of supporting evidence into the final context.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.