BAPO: Behavior-Aware Policy Optimization for Fine-Grained Credit Assignment in Long-Horizon Agents
Abstract
Reinforcement learning (RL) has become an effective approach for improving the sequential decision-making capabilities of large language model (LLM) agents. While outcome-based methods such as Group Relative Policy Optimization (GRPO) offer a simple way to optimize agent policies using task-level rewards, their trajectory-level supervision often fails to distinguish the contributions of intermediate interactions. Recent methods therefore exploit finer-grained structures within agent trajectories to improve credit assignment. A key challenge is to distinguish intermediate states that reflect meaningful task progress from those that contribute little to the final outcome. Changes in an agent's tool-use behavior can provide useful signals for locating such task-relevant states along the trajectory. Building on this insight, we propose **Behavior-Aware Policy Optimization** (BAPO), a simple method for identifying task-relevant intermediate states and using them to construct fine-grained learning signals for policy optimization. BAPO first detects coarse changes in an agent's tool-use behavior to localize candidate transition regions. It then attributes task-level outcomes to these candidate transitions by contrasting behavioral patterns across successful and failed trajectories, identifying transitions associated with task success. Based on this attribution, BAPO redistributes trajectory-level advantages toward actions associated with reaching these states. In this way, BAPO uses task outcomes to provide differentiated supervision for intermediate interactions. Experiments on ALFWorld and Terminal-Bench with Qwen-family models show that BAPO consistently improves agent performance, with success-rate gains of up to 4.5 %. Beyond policy optimization, we find that these transition signals can also guide inference-time trajectory compression, suggesting that they capture behaviorally meaningful structure useful for both training and inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.