Latent Future Goal Representations
Abstract
Goals in goal-conditioned reinforcement learning are usually single observations, yet abstract tasks are more naturally stated as sets of outcomes, for example reaching either of two states. Such abstract goals can be represented in a latent goal space, but critic and policy have no direct way to use this structure. Typically, they must learn again from value targets what those representations mean, so that an agent trained on single goals, for example, does not generalize to more abstract set-valued goals. We introduce Latent Future Goal Representations (LFG), which learns a shared goal space without critic gradients and trains an action-conditioned predictor to map state-action pairs directly into it. Each prediction represents a latent goal region, and a learned bilinear containment relation scores whether a desired goal belongs to that region. These containment scores serve directly as the critic for policy learning. The predictor is trained on observed future goals, with targets discounted by the time until their arrival. In experiments we find that our shared representation enables the critic to inherit the learned goal space structure, allowing generalization to abstract goals. On OGBench with single observation goals, LFG achieves competitive performance across navigation and manipulation with both state and image observations. Ablations show that relational supervision improves control over temporal supervision alone, while the learned containment relation provides an effective critic readout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.