acceptodds
Under review as a conference paper at ICLR 2027

FocusWAM: Adaptive Multi-View Future Prediction for World Action Models

Abstract

World action models (WAMs) use future prediction to guide robot manipulation, but the views needed to anticipate an interaction vary during task execution. Global prediction can miss fine-grained gripper–object dynamics, while predicting all views at every step incurs additional computational cost even when local details contribute little to control. We introduce FocusWAM, a world action model that uses global future features to select when to predict wrist views. It first anticipates scene-level dynamics and determines whether to activate wrist-view prediction for the upcoming action chunk. When activated, the local stream predicts wrist-view dynamics conditioned on global features. The resulting local features refine the global future representation. We supervise the refined representation with a global future prediction objective and use the global and local features to condition action generation. FocusWAM achieves 55.0% average success on RoboCasa GR1 and 50.40% on the randomized split of RoboTwin 2.0 after training on clean demonstrations only, together with 76.50% across four real-world tasks. Across the clean and randomized RoboTwin evaluations, FocusWAM improves average success over an independently trained always-on local policy from 65.0% to 68.3% while achieving a 1.50 inference speedup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.