acceptodds
Under review as a conference paper at ICLR 2027

R3PO: Unified Policy Optimization For Agentic RAG

Abstract

Agentic retrieval-augmented generation (RAG) systems depend on a reasoner and a retriever, yet a schism in their optimization—reinforcement learning (RL) for the reasoner versus indirect signals for the retriever—hinders true end-to-end alignment with task outcomes. We introduce Reasoner-Retriever Reinforcement Policy Optimization (R3PO), a staged framework that applies a common, outcome-driven policy optimization interface to both reasoner and retriever actions. Our approach is built on three pillars: principled policy construction, which defines a stochastic policy that respects the geometry of the pretrained retriever's native scoring function; principled adaptation, which uses a unified RL objective to refine both components with trajectory-level rewards; and principled generalization, which confronts the critical “candidate-to-corpus chasm.” We prove that adapting a retriever on a small candidate pool makes it blind to score shifts that are catastrophic for full-corpus retrieval. R3PO solves this by introducing background documents as anchors during training, grounding the learned policy and ensuring it generalizes beyond the candidate set. This preserves the retriever's efficient inner-product structure, enabling direct, full-corpus deployment. Across seven QA benchmarks and three backbones, R3PO consistently achieves higher accuracy with greater efficiency, delivering relative gains of 2.2%-3.5% in average exact match (EM) while even enhancing an already agent-adapted retriever. These results demonstrate that a principled, unified RL approach unlocks significant performance gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.