acceptodds
Under review as a conference paper at ICLR 2027

Next-WAM: Any-to-any World Action Model for Robotic Manipulation

Abstract

Vision-Language-Action (VLA) models and World-Action Models (WAMs) have advanced robot manipulation by leveraging multimodal understanding and predictive modeling, respectively. Nevertheless, both paradigms struggle to jointly reason over task intent, forecast visual outcomes, and output precise actions. In this work, we introduce Next-WAM, an any-to-any World-Action Model that unifies these capabilities within a single multimodal framework. Built upon a Mixture-of-Transformers (MoT) architecture, our model consists of understanding, generation, and action experts with disjoint parameters and structured joint attention. It fuses modality-specific encodings while enabling cross-modal information exchange. Powered by this design, Next-WAM consumes visual history, produces subtask text as control guidance, predicts future observations, and outputs continuous action chunks. Notably, by masking future-image latents from the action attention stream, Next-WAM leverages visual prediction to supervise control representations, without needing to generate future images at action inference time. Trained without extra embodied pretraining, Next-WAM attains 98.4% success on LIBERO and 70.5% on LIBERO-Plus, outperforming Fast-WAM by 0.8% and 19.0%, respectively. Real-robot experiments further validate Next-WAM’s capacity to reason from task instructions for manipulation. Code will be made public.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.