acceptodds
Under review as a conference paper at ICLR 2027

GCHR: Goal-Conditioned Hindsight Regularization via Compositional Priors

Abstract

In goal-conditioned reinforcement learning (GCRL) with sparse rewards, the agent receives little learning signal. Hindsight experience replay (HER) relabels trajectories with the goals they achieved, but it treats each relabeled transition independently. The actor therefore receives no direct guidance from the structure of whole trajectories. We propose GCHR (Goal-Conditioned Hindsight Regularization), a backbone-agnostic framework that distills past trajectories into hindsight compositional priors and regularizes the actor toward them. A hindsight behaviour prior anchors the policy to relabeled actions. A hindsight goal prior aggregates the target policy's behaviours toward waypoints on the same trajectory to transfer knowledge across goals. Both priors come from the replay buffer and a target copy of the actor, so GCHR trains no additional network. Theoretically, the goal prior extends action coverage beyond self-imitation, and its training loss upper-bounds the divergence to the exact prior in expectation. Under idealized conditions, a via-goal value that measures cross-goal transfer does not degrade as the policy improves. Experiments on three benchmark suites with four backbones (DDPG, SAC, TD3, CRL) and state or visual observations show gains in sample efficiency and final performance on most tasks. Ablation, goal-corruption, and coverage studies link these gains to trajectory-aligned cross-goal transfer.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.