MoAct: Learning Visual Transition Dynamics for Robot Action Generation
Abstract
Modeling future observations allows vision-language-action (VLA) models to anticipate scene changes and improve action generation. However, extracting latent representations that retain the information needed for accurate prediction and control remains a key challenge. To address this challenge, we introduce MoAct, a framework that uses compact visual-transition tokens to model changes in future visual context and guide action generation. MoAct learns visual-transition dynamics from large-scale video data and connects a visual-transition dynamics expert with an action expert through a dual-stream Mixture-of-Transformers (MoT) architecture. This design allows predicted scene changes to guide policy learning. Extensive experiments in simulation and the real world show that MoAct outperforms existing future-guided methods. Egocentric human videos also help MoAct retain much of its performance with fewer robot demonstrations. For scene adaptation, MoAct requires only additional egocentric human videos from the target scenes, with no new robot demonstrations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.