acceptodds
Under review as a conference paper at ICLR 2027

World-Gymnast: Training Robots with Reinforcement Learning in a World Model

Abstract

Robot learning is bottlenecked by costly real-world interaction. Supervised finetuning (SFT) is limited by expert demonstration coverage, while reinforcement learning (RL) in software simulators faces a sim-to-real gap in manipulation. We ask whether training policies in world models learned from real-world video-action data can improve real-robot performance more effectively than either alternative. We propose World-Gymnast, a framework for improving vision-language-action (VLA) policies through reinforcement learning in a pretrained world model. Given an initial image and a language instruction, the policy learns through imagined interactions, with a zero-shot vision-language model (VLM) judging task completion. We show that a single policy trained across multiple tasks in this way improves performance on real-robot tasks and environments held out from RL training. On the Bridge robot setup, World-Gymnast achieves up to 18× the success rate of supervised fine-tuning and 2× that of RL in a software simulator. More importantly, World-Gymnast demonstrates intriguing capabilities of RL with a world model, including training on diverse language instructions and novel scenes from the world model, test-time training in a novel scene, and online iterative world model and policy improvement. Our results suggest learning a world model and training robot policies in the cloud could be the key to bridging the gap between robots that work in demonstrations and robots that can work in anyone’s household.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.