acceptodds
Under review as a conference paper at ICLR 2027

Learning to Shape the Reward for Agentic Reinforcement Learning via Bilevel Optimization

Abstract

Reinforcement learning for agentic systems often relies on trajectory-level rewards, leading to poor credit assignment in multi-turn tasks. A common approach is to introduce turn-level rewards, such as heuristic signals, LLM judges, or learned rewards, to provide dense feedback. However, existing reward learning approaches typically optimize the reward based on trajectories generated by the current policy, while treating the trajectory distribution as independent of the learned reward. In reality, the learned reward determines how the policy is optimized, which in turn affects the trajectories generated by the policy. Ignoring this dependency results in an incomplete gradient for reward learning. This paper proposes Bilevel Agentic Reward Shaping (BARS), a framework that explicitly captures this reward-policy-trajectory dependency through bilevel optimization. The upper level learns a turn-level reward from outcome feedback, while the lower level optimizes the policy under the learned reward. We develop an efficient algorithm that alternates between reward learning and policy updates, and establish a convergence rate of . Empirical results demonstrate that BARS achieves higher accuracy and improved training stability compared with baseline methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.