DreamTamer: From Lucky Samples to Reliable Futures in Action Conditioned World Models
Abstract
Action-conditioned world models (ACWMs) predict future observations from visual histories and robot actions, and must faithfully capture the consequences of those actions. However, even with identical observations and actions, pretrained ACWMs can produce substantially different interaction outcomes across random seeds, making faithful prediction inconsistent across samples. To understand this variability, we conduct denoising-stage branching experiments and find that early-stage stochasticity strongly affects prediction fidelity and task-critical interaction dynamics, whereas later-stage stochasticity primarily affects visual details. This finding motivates a reinforcement learning post-training framework that concentrates optimization on denoising transitions with greater influence on prediction fidelity. Specifically, we combine timestep-aware gradient correction with variance-aware timestep sampling to direct learning toward these influential transitions. To further improve reliability across samples, we introduce a lower-bound penalty that targets low-fidelity predictions without directly penalizing favorable deviations. Experiments on Ctrl-World, DreamDojo, and Genie Envisioner demonstrate improvements in average prediction quality, the quality of low-fidelity predictions, and consistency across repeated samples. Together, these results show that our framework makes faithful action-conditioned prediction more reliable across random seeds. We will release all training and evaluation code and model checkpoints to support reproducibility and further research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.