acceptodds
Under review as a conference paper at ICLR 2027

Harlo: Hierarchical Agentic RL with Information-Gain Guided Credit Assignment

Abstract

Agentic reinforcement learning (RL) has advanced LLM-based agents, yet most existing algorithms are confined to single-agent settings. The Planner-Executor architecture, where a planner generates high-level plans and an executor executes step-by-step actions, offers a promising paradigm for such tasks, but efficient and stable RL methods for co-optimizing both agents remain lacking. In this paper, we propose Harlo, a hierarchical agentic RL framework that addresses the credit assignment problem through an information-theoretic approach. Our key insight is that information gain—measured as the behavioral influence of the plan on the executor's actions via counterfactual evaluation—serves as a natural signal for allocating credit. This signal enables a two-level credit assignment strategy that dynamically reshapes advantage values for both agents, ensuring stable joint optimization and co-evolution. Evaluated on ALFWorld, ScienceWorld and WebShop, Harlo consistently outperforms strong and diverse baselines. Code is available at https://anonymous.4open.science/r/Harlo-F2F5.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.