Coverage Maximizing Exploration
Abstract
We propose batched goals exploration (BGE), an adaptive exploration algorithm that simultaneously targets multiple novel regions in the state space without any online RL updates. This is in contrast to standard RL exploration algorithms that adapt to new reward information by repeatedly updating the exploration policy online with RL objectives, which can be expensive and unstable. BGE avoids online RL updates by decoupling behavior optimization from high-level online exploration. In particular, BGE first pre-trains a goal-conditioned agent on a prior offline dataset, then online optimizes a batch of goals such that the expected state occupancy achieved by rolling out the pre-trained goal-conditioned agent conditioned on these goals maximizes the coverage over potential high-reward states. Across challenging long-horizon, sparse-reward tasks across locomotion and manipulation, BGE is able to leverage unlabeled prior data to achieve better online sample-efficiency compared to prior methods, while only using a fraction of the computational cost online.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.