Iterative Self-Improvement with World-Expectation Optimization for World-Action Model
Abstract
Robot foundation models have enabled increasingly general manipulation through large-scale pretraining. Recent world action models (WAMs) further unify action generation and future prediction, providing a foundation for robots to learn both how to act and what their actions cause. As robots interact with their environments, each execution provides new evidence about the effectiveness of their actions and the accuracy of their predictions, creating an opportunity for continued self-improvement beyond pretraining. However, existing self-improvement approaches often evolve policies with external predictive models, without directly improving predictive capabilities within the action-generating model itself. We introduce World-ISI, an iterative framework that turns this execution feedback into the joint evolution of action generation, world prediction, and value estimation within a single WAM. To leverage the WAM’s own predictions of future task outcomes during self-improvement, we further propose World-Expectation Optimization (WEO). WEO uses the difference between observed returns and the model’s own value predictions to weight learning from successful actions, while using failed executions in dedicated world- and value-prediction modes and retaining auxiliary prediction losses on successful policy samples. Applied to Cosmos Policy, World-ISI improves average success on 24 RoboCasa tasks from 66.7% to 70.1% after four iterations. In a controlled comparison at 6,000 training steps, WEO achieves 69.0% success versus 67.5% for standard supervised fine-tuning, supporting the effectiveness of learning from the model’s own expectations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.