Latent Successor Models
Abstract
An effective reinforcement learning (RL) agent must reason about both the immediate outcomes and long-term consequences of its actions. This challenge is especially pronounced when learning from pixel observations. Many prior RL methods learn visual representations for control by combining representation learning with reward prediction and value estimation. How can we acquire such representations without directly training them to predict rewards or values? In this paper, we propose *Latent Successor Models* (LSM) an approach that learns actionable visual representations through generative modeling of future outcomes at arbitrary horizons and behavioral cloning. This removes the need for direct reward and value prediction or pixel reconstruction when learning these representations. By sampling distant future outcomes under a policy, we construct targets for value estimation without predicting every intermediate state. Our analysis shows that directly predicting future outcomes yields more accurate value estimates at longer horizons. Experiments across a range of visual offline RL benchmarks ( tasks) demonstrate substantial improvements over prior model-free and model-based approaches, with particularly large gains on challenging long-horizon tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.