P-WAM: Perception-World-Action Models
Abstract
World Action Models (WAMs) have recently shown strong potential for robot manipulation by jointly modeling future visual states and actions, thereby capturing environment dynamics and physical regularities for closed-loop control. However, existing WAMs typically rely on frozen reconstruction-oriented visual VAEs for initial environment perception, which may limit their ability to capture object-level semantics and fine-grained spatial information. In this work, we propose Perception-World-Action Models (P-WAM), a simple and general framework that enhances WAMs with pretrained visual perception encoders while preserving their world dynamics modeling. P-WAM introduces complementary semantic and spatial perceptual features from discriminative visual encoders into the video-action DiT through lightweight connector modules. We instantiate P-WAM with SigLIP and DINOv3 to provide language-aligned semantics and dense spatial representations. Through comprehensive representation probing, experiments on the challenging RoboCasa simulation benchmark, and diverse real-world robot evaluations, we show that P-WAM improves the object-level semantic and fine-grained spatial information encoded in video-action DiT representations and leads to stronger closed-loop policy performance. These results suggest that integrating visual perception with generative world-action modeling is a promising direction for improving WAMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.