Hierarchical Reflective Policy Optimization for LongHorizon Agentic Reinforcement Learning
Abstract
Agentic reinforcement learning (RL) has emerged as the core paradigm for empowering large language model (LLM) agents to accomplish open-ended, long-horizon interactive tasks. However, existing agentic RL methods suffer from two critical bottlenecks that severely limit real-world deployment: coarse trajectory-level credit assignment that fails to capture step-wise valid reasoning in failed rollouts, and static policy optimization that cannot adapt the exploration–exploitation trade-off to dynamic task complexity. To address these fundamental limitations, we propose Hierarchical Reflective Policy Optimization (Hrpo), a novel agentic RL framework equipped with dual-level reflective credit assignment and adaptive curriculum exploration for long-horizon agent decision-making. Specifically, we first construct a trajectory state-transition graph to model hierarchical agent–environment interaction dynamics, which decomposes holistic episode rewards into finegrained step-level and sub-task-level credit signals, effectively excavating latent valuable reasoning steps obscured in incomplete or failed trajectories. Furthermore, we design a task-aware reflective exploration module that enables the agent to self-summarize historical interaction errors and successful experiences during rollouts, dynamically adjusting policy exploration intensity and avoiding redundant trial-and-error in repeated scenarios. Unlike conventional fixed-strategy RL and static agent refinement methods, Hrpo unifies experiential reflection, hierarchical credit assignment and adaptive policy update in a single end-to-end optimization pipeline, significantly improving the stability and efficiency of agent training. On a controlled long-horizon agent evaluation suite, Hrpo reaches an 80% task success rate using 88% fewer environment interactions than the group-relative baseline and attributes failures to the true faulty sub-task with 72.8% top-1 accuracy versus 12.4% for trajectory-level credit under fault injection. Ablations attribute these gains primarily to the hierarchical credit assignment; the reflective exploration component yields no measurable additional benefit in the small-scale suite—a negative result we report transparently and revisit under the large-benchmark protocol.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.