Iso-WAM: Rethinking Representation Alignment in World Action Models
Abstract
World-Action Models (WAMs) adapt pretrained video generators for robot manipulation, but pixel-level future prediction entangles action-induced visual changes with scene appearances unrelated to the action, making control brittle under visual distribution shifts. Aligning policies with future static representation, such as semantic or geometric features, offers an alternative source of supervision, but these representations may still retain irrelevant background information. We propose Iso-WAM, which disentangles physical dynamics from visual appearance through dynamic representation alignment. Iso-WAM uses learnable action queries to predict latent action targets provided by a pretrained Unified Latent Action Model (Uni-LAM). Uni-LAM integrates semantic and geometric transitions into a single latent action space for alignment, providing the policy with comprehensive and robust dynamic guidance. Experiments show that Iso-WAM achieves strong in-distribution performance on the standard LIBERO benchmark and robust zero-shot OOD generalization on LIBERO-Plus and real-world robot deployments. In particular, Iso-WAM attains a 75.1% average success rate on LIBERO-Plus, improving over the Fast-WAM-Joint baseline by 7.5%. Code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.