FocusWAM: Adaptive Multi-View Future Prediction for World Action Models
Abstract
World action models (WAMs) use future prediction to guide robot manipulation, but the views needed to anticipate an interaction vary during task execution. Global prediction can miss fine-grained gripper–object dynamics, while predicting all views at every step incurs additional computational cost even when local details contribute little to control. We introduce FocusWAM, a world action model that uses global future features to select when to predict wrist views. It first anticipates scene-level dynamics and determines whether to activate wrist-view prediction for the upcoming action chunk. When activated, the local stream predicts wrist-view dynamics conditioned on global features. The resulting local features refine the global future representation. We supervise the refined representation with a global future prediction objective and use the global and local features to condition action generation. FocusWAM achieves 55.0% average success on RoboCasa GR1 and 50.40% on the randomized split of RoboTwin 2.0 after training on clean demonstrations only, together with 76.50% across four real-world tasks. Across the clean and randomized RoboTwin evaluations, FocusWAM improves average success over an independently trained always-on local policy from 65.0% to 68.3% while achieving a 1.50 inference speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.