acceptodds
Under review as a conference paper at ICLR 2027

Model-Based Meta-Learning for Algorithm Discovery

Abstract

Reinforcement Learning (RL) algorithms are typically hand-crafted through a slow and iterative scientific process. While meta-RL promises to automate RL algorithm discovery, high simulation costs have limited research. In this work, we introduce Model-Based Meta-Learning (MBML), a novel approach that uses learned world models as an efficient alternative to environment simulation. Policies trained in these world models fail to transfer to the simulator, but algorithms discovered in world models transfer successfully. We show that MBML matches the performance of meta-RL with environment simulators at a fraction of the time and compute. By substantially reducing the computational cost of meta-training, MBML lowers the barrier to entry for meta-RL research while enabling algorithm discovery at a larger scale. We demonstrate this by using MBML to meta-learn a novel algorithm that significantly outperforms prior work on unseen continuous control environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.