acceptodds
Under review as a conference paper at ICLR 2027

Hierarchical Contrastive Control: Learning Subgoals via Future-State Density Estimation

Abstract

Hierarchical reinforcement learning (HRL) provides a principled framework for solving long-horizon decision-making tasks by jointly learning a high-level policy that predicts subgoals and a low-level policy that executes primitive actions. However, HRL suffers from training instability because the high-level transition dynamics and rewards change as the low-level policy improves, leading to unstable higher-level temporal-difference learning. We investigate an alternative formulation of hierarchical goal-conditioned RL based on future-state density estimation via contrastive learning, where the high-level policy predicts subgoals that maximize the induced future-state density assigned to the final goal, instead of relying on higher-level temporal-difference bootstrapping. However, directly applying this formulation to hierarchical control is non-trivial, as the high-level policy may predict infeasible subgoals for the current low-level policy, thus hindering final-goal reachability. To address this, we propose Hierarchical Contrastive Control (HCC), a constrained optimization approach that regularizes the future-state density objective with a lower-level optimality constraint that guides the high-level policy toward feasible subgoal predictions for the current low-level policy. We perform experiments on sparse-reward robotic tasks, showing that HCC achieves over 55% success rates, outperforms state-of-the-art baselines and yields feasible subgoals, establishing hierarchical future-state density estimation as an effective learning signal for high-level policy optimization in jointly trained hierarchical control.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.