acceptodds
Under review as a conference paper at ICLR 2027

Recursively Self-Improving World Models for Reinforcement Learning

Abstract

Model-based reinforcement learning (MBRL) achieves strong data efficiency by leveraging synthetic experience for policy optimization; however, training a world model still requires substantial real interaction data, not only for accurate predictions but also for adequate coverage of real transitions. In this work, we investigate whether a world model can improve itself through its own generated experience to better support policy learning. We show that self-training amplifies concentration tendencies in generated experience, arising from mode-seeking sampling in unconditional generators and coupled model-policy interaction in conditional dynamics models. Reversing the resulting parameter displacement through negative extrapolation steers the world model toward generating more diverse and useful synthetic experience for downstream policy learning. Across both unconditional and conditional world-model formulations, our approach improves policy performance and data efficiency over the respective backbones on MuJoCo, DeepMind Control Suite, and Meta-World tasks with modest additional computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.