acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Reflective Policy Optimization for LongHorizon Agentic Reinforcement Learning

Abstract

Agentic reinforcement learning (RL) has emerged as the core paradigm for empowering large language model (LLM) agents to accomplish open-ended, long-horizon interactive tasks. However, existing agentic RL methods suffer from two critical bottlenecks that severely limit real-world deployment: coarse trajectory-level credit assignment that fails to capture step-wise valid reasoning in failed rollouts, and static policy optimization that cannot adapt the exploration–exploitation trade-off to dynamic task complexity. To address these fundamental limitations, we propose Hierarchical Reflective Policy Optimization (Hrpo), a novel agentic RL framework equipped with dual-level reflective credit assignment and adaptive curriculum exploration for long-horizon agent decision-making. Specifically, we first construct a trajectory state-transition graph to model hierarchical agent–environment interaction dynamics, which decomposes holistic episode rewards into finegrained step-level and sub-task-level credit signals, effectively excavating latent valuable reasoning steps obscured in incomplete or failed trajectories. Furthermore, we design a task-aware reflective exploration module that enables the agent to self-summarize historical interaction errors and successful experiences during rollouts, dynamically adjusting policy exploration intensity and avoiding redundant trial-and-error in repeated scenarios. Unlike conventional fixed-strategy RL and static agent refinement methods, Hrpo unifies experiential reflection, hierarchical credit assignment and adaptive policy update in a single end-to-end optimization pipeline, significantly improving the stability and efficiency of agent training. On a controlled long-horizon agent evaluation suite, Hrpo reaches an 80% task success rate using 88% fewer environment interactions than the group-relative baseline and attributes failures to the true faulty sub-task with 72.8% top-1 accuracy versus 12.4% for trajectory-level credit under fault injection. Ablations attribute these gains primarily to the hierarchical credit assignment; the reflective exploration component yields no measurable additional benefit in the small-scale suite—a negative result we report transparently and revisit under the large-benchmark protocol.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.