SIP-WAM: Enhancing World Action Models with Interaction-aware Spatiotemporal Prediction
Abstract
World Action Models (WAMs) have demonstrated strong potential for robot policy learning by jointly modeling future visual observations and actions. However, existing RGB-centric WAMs remain limited in learning interaction-aware spatiotemporal representations. Temporally sparse visual prediction provides insufficient supervision for fine-grained interaction dynamics, whereas temporally dense RGB prediction incurs substantial computational overhead. Meanwhile, full-frame RGB objectives do not explicitly prioritize task-relevant interaction regions, limiting interaction-aware spatial grounding. To address these limitations, we introduce SIP-WAM, a framework that enhances WAMs with Interaction-aware Spatiotemporal Prediction. Specifically, alongside the WAM backbone, an Interaction DiT jointly predicts the current spatial locations and temporally dense future trajectories of a sparse set of points uniformly sampled on interaction participants, providing efficient and control-relevant interaction supervision for both spatial grounding and fine-grained temporal dynamics. To avoid additional inference overhead, we couple the two branches through layerwise asymmetric joint attention, enabling the Interaction DiT to guide representation learning during training and be removed at inference. Extensive experiments on RoboCasa, LIBERO-Plus, and real-world single-arm and bimanual manipulation demonstrate strong policy performance and generalization, validating the effectiveness of interaction-aware spatiotemporal supervision for world action modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.