Deep Exploration Through Pre-image State Distribution Matching (PIM)
Abstract
We present an algorithm for reinforcement learning that explores and learns to solve long-horizon, goal-reaching tasks without any environment rewards, demonstrations, curriculum, or prior knowledge. With our approach, a simulated robot can learn to construct complex 8-cube configurations, a task not previously solved by prior methods based on intrinsic motivation or subgoal selection. We show that the generalization capabilities of generative models trained from scratch can drive exploration, even in difficult, high-dimensional tasks. We train a generative model with only the online data collected by the policy to predict what states are likely to precede a goal state. Before the goal state been reached, this process is analogous to prompting an image generation model on an unseen prompt. Initially, this generative model makes poor predictions, but when the policy tries and fails to visit those states, it collects the data required to fix the model's predictions errors. Our algorithm outperforms a diverse set of baselines across multiple long-horizon tasks involving robotic manipulation, construction, and navigation. On the AllegroHand dexterous grasp-and-throw task, even the best-performing baseline does not obtain reliable success despite using more interactions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.