acceptodds
Under review as a conference paper at ICLR 2027

HERAKLES: Hierarchical Skill Compilation for LLM Agents in Compositional Goal Spaces

Abstract

LLM agents fine-tuned with reinforcement learning exploit common-sense knowledge to act in interactive environments, but typically decide every elementary action. In compositional goal spaces, where reaching a goal requires reaching others first, this forces credit assignment over horizons that grow with goal depth. This in turn leads to considerable decrease of sample efficiency as goals get more complex. We posit that the compositionality of behaviours follows that of their linguistic descriptions, so that an LLM is best used to reason over goals, while low-level routines run without language. Existing hierarchies built on this division rely on skills supplied by humans, or grow them without direction. We introduce HERAKLES, in which an LLM planner and a lightweight controller acting from raw observations are trained concurrently by RL, without any pre-trained skill library. The goal space is given, but its prerequisite structure is not: the planner selects intermediate goals among those the controller is estimated to master, and trajectories reaching a goal are compiled into the controller, turning mastered goals into options for harder ones and amortizing past learning. On Crafter, HERAKLES keeps improving where flat and hierarchical baselines plateau, and degrades markedly less on compositional variants of known goals.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.