What goals are ahead? Kerl: Kernel Mean Embedding to scale RL
Abstract
Goal-conditioned RL commonly uses hindsight relabeling to associate each state-action pair in a replay buffer with a goal reached later in the same trajectory. By mixing experience collected under different commanded goals, this implicitly samples from a distribution over goals that is rarely made explicit. We formalize it as the posterior-averaged goal occupancy measure and introduce Kerl, an actor-critic method grounded in its kernel mean embedding. With a characteristic kernel, the kernel mean embedding uniquely determines the underlying distribution. In practice, Kerl approximates it in finite dimensions by training the critic to predict the expected random Fourier features of future goals. These fixed features replace the learned goal encoder of contrastive RL, while a squared-error regression loss replaces its contrastive objective, reducing computational cost within the same actor-critic structure. The learned representation enables reconstruction of the goal distribution, revealing which goals are expected to follow a given state-action pair. Our experiments and ablation studies show that Kerl scales effectively with network depth, achieving large gains on the hardest JaxGCRL tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.