AsyncMoT: Interaction-Guided Asynchronous Mixture of Transformers for Generalized and Robust Robotic Manipulation
Abstract
Robotic manipulation involves interaction contexts that persist across control steps but change at irregular, interaction-dependent moments. Yet existing policies typically couple interaction reasoning to control frequency or predefined task structure, creating a mismatch between when robots act and when they need to reason. We introduce AsyncMoT, an interaction-guided asynchronous Mixture of Transformers that makes when to reason an explicit policy decision. Given a high-level task instruction, AsyncMoT maintains a structured interaction state capturing the current stage, target object, and action-relevant region, and selectively refreshes it as the interaction evolves. Between updates, cached Key-Value features preserve interaction guidance across control steps, while the action expert continuously adapts motor commands to current visual and proprioceptive observations. This decoupling further enables interaction supervision to be shared across heterogeneous robot tasks and human videos, without requiring human-to-robot action correspondence. AsyncMoT achieves average success rates of 98.1% on LIBERO and 88.9% on LIBERO-Plus, 66.3% on RoboCasa-24 and 35.9% on RoboCasa365, demonstrating strong performance across long-horizon, distribution-shifted, and compositional manipulation settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.