Model-Based Meta-Learning for Algorithm Discovery
Abstract
Reinforcement Learning (RL) algorithms are typically hand-crafted through a slow and iterative scientific process. While meta-RL promises to automate RL algorithm discovery, high simulation costs have limited research. In this work, we introduce Model-Based Meta-Learning (MBML), a novel approach that uses learned world models as an efficient alternative to environment simulation. Policies trained in these world models fail to transfer to the simulator, but algorithms discovered in world models transfer successfully. We show that MBML matches the performance of meta-RL with environment simulators at a fraction of the time and compute. By substantially reducing the computational cost of meta-training, MBML lowers the barrier to entry for meta-RL research while enabling algorithm discovery at a larger scale. We demonstrate this by using MBML to meta-learn a novel algorithm that significantly outperforms prior work on unseen continuous control environments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.