GAGE: Training General Agents with Go-Explore Demonstrations
Abstract
Go-Explore is a powerful family of methods for solving hard-exploration problems. These approaches learn by first returning to and then exploring from promising states discovered during training. However, despite dramatically outperforming other exploration techniques in video game, robotics, and text-based reasoning tasks, prior work has not used Go-Explore to train policies that generalize across problem instances. We introduce General Agents via Go-Explore (GAGE), a new method for training general policies in highly variable, sparse-reward environments. GAGE searches for successful trajectories on individual task instances with Go-Explore, distills these demonstrated behaviors into a policy with supervised learning, and then fine-tunes the policy with reinforcement learning. Our experiments on Procgen and MiniHack, two procedurally generated reinforcement learning benchmarks, demonstrate that GAGE generalizes to unseen problem instances while standard reinforcement learning methods struggle. We show that GAGE can match or exceed the sample efficiency of popular RL methods, that its performance reliably improves with additional samples, and that GAGE is robust to increasing reward sparsity. Finally, we show that GAGE can train a single multi-task policy that performs well across several environments. Our work suggests that combining Go-Explore with supervised learning is a promising and underexplored paradigm for training general agents in hard-exploration domains.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.