R3PO: Unified Policy Optimization For Agentic RAG
Abstract
Agentic retrieval-augmented generation (RAG) systems depend on a reasoner and a retriever, yet a schism in their optimization—reinforcement learning (RL) for the reasoner versus indirect signals for the retriever—hinders true end-to-end alignment with task outcomes. We introduce Reasoner-Retriever Reinforcement Policy Optimization (R3PO), a staged framework that applies a common, outcome-driven policy optimization interface to both reasoner and retriever actions. Our approach is built on three pillars: principled policy construction, which defines a stochastic policy that respects the geometry of the pretrained retriever's native scoring function; principled adaptation, which uses a unified RL objective to refine both components with trajectory-level rewards; and principled generalization, which confronts the critical “candidate-to-corpus chasm.” We prove that adapting a retriever on a small candidate pool makes it blind to score shifts that are catastrophic for full-corpus retrieval. R3PO solves this by introducing background documents as anchors during training, grounding the learned policy and ensuring it generalizes beyond the candidate set. This preserves the retriever's efficient inner-product structure, enabling direct, full-corpus deployment. Across seven QA benchmarks and three backbones, R3PO consistently achieves higher accuracy with greater efficiency, delivering relative gains of 2.2%-3.5% in average exact match (EM) while even enhancing an already agent-adapted retriever. These results demonstrate that a principled, unified RL approach unlocks significant performance gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.