acceptodds
Under review as a conference paper at ICLR 2027

Jumpy Planning over Policy Occupancies Improves Online Goal-Conditioned RL

Abstract

Online goal-conditioned reinforcement learning aims to learn a single policy for many goals, but sparse feedback makes long-horizon behaviors difficult to discover from interaction alone. Planning with temporal abstractions through Geometric Policy Composition could accelerate the acquisition of such behaviors but existing methods have largely focused on fixed, pretrained policies and world models. We introduce Online Geometric Policy Composition (OGPC), which plans over short subgoal sequences to guide exploration while learning both the policy and its world model online from scratch. OGPC uses a Geometric Horizon Model to predict which states the policy reaches while pursuing each subgoal, and chains these predictions to evaluate candidate sequences. To score candidate plans reliably under sparse rewards, we replace the final sampled reward in the existing Monte Carlo estimator with a learned estimate of its conditional success probability. With this estimator, OGPC improves sample efficiency and final performance over a strong, depth-scaled Contrastive Reinforcement Learning policy across nine JaxGCRL tasks, including high-dimensional Humanoid environments. The gains persist without planning at evaluation, showing that temporally abstract planning improves the learned policy rather than only decision-time behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.