acceptodds
Under review as a conference paper at ICLR 2027

Stable Temporal Abstraction for Hierarchical Reinforcement Learning in Language Model Agents

Abstract

Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (), a general constrained boundary-policy framework that represents premature replanning and stale persistence as constraint costs. We instantiate in two forms: uses duration-based proxies for premature replanning and stale persistence, providing a simple environment-agnostic realization; learns semantic completion probabilities from environment-grounded supervision and converts them into state-dependent boundary costs. Both instantiations apply their Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, original critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two primary backbones and two benchmarks, improves success over a strong hierarchical baseline by and points on ALFWorld and WebShop with Qwen3-0.6B and by and points with Llama-3.2-1B-Instruct. A separate Qwen2.5-1.5B-Instruct ALFWorld replication reaches peak success with versus for the released HiPER host. The code is available at https://anonymous.4open.science/r/STAC-EFF9.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.