acceptodds
Under review as a conference paper at ICLR 2027

ANGULAR REACHABILITY REPRESENTATIONS FOR OFFLINE GOAL-CONDITIONED REINFORCEMENT LEARNING

Abstract

Offline goal-conditioned reinforcement learning (GCRL) provides a promising framework for learning generalizable goal-reaching policies from large-scale offline trajectory datasets. However, existing approaches often suffer from inadequate modeling of state-goal relationships, leading to suboptimal policy learning and limited deployment accuracy. In this paper, we introduce a novel contrastive representation learning framework for offline GCRL based on unit hyperspherical state-goal embeddings. Our approach learns structured representations of states and goals on two unit hyperspheres, where the reachability between a state and a goal is characterized by their angular relationship. Specifically, this reachability measure can be interpreted as a goal-conditioned value function defined by the future-state occupancy measure, and is equivalently represented by the cosine similarity between the corresponding unit embeddings. We refer to these embeddings as Angular Reachability Representations (ARR). Building upon the learned representations, we propose a novel offline reinforcement learning algorithm that performs policy optimization based on ARR, improving data utilization and mitigating distribution shift. Furthermore, we introduce a test-time planning strategy that leverages ARR to refine action selection during deployment without additional training. Experiments on OGBench demonstrate the effectiveness of the proposed representation, policy-learning method, and test-time planner across offline goal-conditioned control tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.