acceptodds
Under review as a conference paper at ICLR 2027

SWAM: Self-Supervised World Action Modeling via Asymmetric Denoising

Abstract

We revisit denoising in world action models from a self-supervised representation learning perspective, shifting emphasis away from video generation. Rather than relying on costly, error-prone future-video rollouts, robot policies need robust visual features that support action prediction from current observations. We therefore introduce the Self-Supervised World Action Model (), which trains action-relevant robust visual features through asymmetric denoising. uses Denoising Autoencoding (DAE) on current observations with a noise schedule less aggressive than that of video generation, providing auxiliary self-supervision for the visual backbone. For future observations, Masked Autoencoding (MAE) motivates extreme-noise denoising that limits access to the target and enables single-pass visual inference. achieves success rates of 99.1% on LIBERO and 83.4% on LIBERO-Plus. On LIBERO-Plus, it exceeds Fast-WAM and Fast-WAM-Joint by 31.9 and 15.0 percentage points, respectively, with comparable inference-path latency to Fast-WAM and over 50% lower latency than Fast-WAM-Joint. These results support asymmetric denoising as an effective self-supervised objective for robot policy learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.