acceptodds
Under review as a conference paper at ICLR 2027

DreamEvolve: Generalizable Robot Policy Self-Evolution through Latent World Models

Abstract

Generalist robot policies have shown impressive manipulation skills, but deploying them reliably in new scenes and tasks remains difficult. Real-world evaluation and improvement require extensive rollouts and human demonstrations, both of which are slow and difficult to scale. World models offer a scalable alternative by allowing policies to roll out in imagination, yet existing pixel-level models entangle task dynamics with appearance and struggle to provide reliable training signals across diverse scenes and novel tasks. We present DreamEvolve, a robotic self-evolution framework built on a controllable latent world model. A Joint-Attention-of-Mixed-Experts architecture learns action-conditioned latent dynamics, while Latent Progress Reward Optimization efficiently converts imagined rollout groups into group-relative progress signals. Across simulation and real-world experiments, DreamEvolve substantially improves policy robustness under out-of-distribution shifts, including cluttered scenes, long-horizon compositional tasks, and challenging deployment conditions. On LIBERO-Pro, DreamEvolve achieves strong cross-scene, cross-task, and out-of-distribution generalization after a single round of imagination-based post-training. DreamEvolve also raises the average success rate under real-world perturbations from 10.1% to 61.7%. We also open-source DreamEvolveBench, the first large-scale paired scene-variation dataset for controllable robotic world models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.