Quasimetric Distance with Flow-Based Planning for Offline Hierarchical Goal-Conditioned RL
Abstract
Offline hierarchical goal-conditioned reinforcement learning (GCRL) decomposes long-horizon control into high-level subgoal selection and low-level execution. However, existing temporal-difference (TD) value-based methods often yield unstable high-level subgoal rankings, because they derive advantages from discounted value differences where long-range progress signals are heavily compressed and easily overwhelmed by bootstrapped approximation errors. By contrast, we revisit hierarchical GCRL through the lens of the value-quasimetric distance equivalence, and show theoretically and empirically that an undiscounted quasimetric distance induces a long-horizon-consistent ordering over subgoals, while its triangle-inequality geometry naturally supports stitching trajectory fragments. Nevertheless, preserving such equivalence requires high-level planning to choose subgoals directly in the original state space, thereby aligning policy optimization with value evaluation. This reformulates high-level policy learning into multimodal density modeling over feasible intermediate states induced by different routes and stitching choices. Motivated by these insights, we propose Hierarchical offline goal-conditioned RL with shared Quasimetric distance and conditional normalizing Flows (HiQFlow), which uses a ranking order-stable quasimetric distance across both hierarchy levels and a conditional normalizing flow to model multimodal subgoals with tractable likelihoods and efficient sampling. Evaluations on OGBench show that HiQFlow outperforms strong offline GCRL baselines, especially on long-horizon, stitching, and exploration-heavy tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.