Entity-Centric Exploration through Latent Dynamics for Visual Reinforcement Learning
Abstract
Exploration is central to reinforcement learning, but it remains challenging in long-horizon tasks with high-dimensional visual inputs. In many such tasks, important structure is not captured at the level of raw pixels or full states, but at the level of entities and their interactions. In this work, we explore how entity structure can be leveraged to make exploration more efficient under visual observations. Through a theoretical connection between dynamics-prediction error and count-based exploration via information gain, we show that exploiting entity structure can reduce exploration complexity, which motivates our entity-centric approach to exploration. Inspired by this, we introduce ECLAIR, an entity-centric exploration method that learns unsupervised, entity-centric representations online and uses a latent dynamics model to estimate entity-level novelty. By training online, ECLAIR allows the entity abstraction to improve as the agent encounters new states. We evaluate ECLAIR on challenging sparse-reward visual RL tasks and demonstrate that its exploration efficiency outperforms prior exploration baselines across skill-based navigation, dynamic control, and robotic manipulation. Additional visualizations and video rollouts are available on our anonymous website: https://sites.google.com/view/eclair-rl.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.