Latent-Predictive Universal Horizon Models for Long Horizon Zero-Shot RL
Abstract
While behavior foundation models (BFMs) have achieved impressive zero-shot performance on various human locomotion tasks, these methods struggle on tasks that require longer horizon reasoning. Recent work has explored using geometric horizon models (GHMs) for planning on top of pre-trained BFMs to achieve longer horizon tasks; however, such methods require significant test-time compute, operate in raw observation space, and require a complex two-stage training process. In this work, we instead explore training horizon models in latent space and as an auxiliary loss during BFM training. In particular, we propose a novel latent-predictive Universal Horizon Model, which we learn jointly with a latent-predictive BFM with both models shaping the state- and task-encoding spaces. We demonstrate that this auxiliary representation learning objective is theoretically compatible with the primary TD-JEPA objective, with respect to the learned state and task encodings, and we demonstrate empirically that it can be beneficial for downstream zero-shot performance. Finally, we show that this UHM model can be used to perform model-based updates in the zero-shot RL setting, allowing us to stably train a latent-predictive BFM at discounts as high as .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.