Actions as Rays: Repurposing Video Diffusion Models for Robot Control with Visual Action Representations
Abstract
Pretrained video diffusion models have learned rich priors over visual dynamics and physical interactions, yet adapting them to robot manipulation remains challenging, as existing approaches often rely on dedicated action heads, specialized tokenizers, or auxiliary policy networks to translate between visual observations and low-dimensional action signals. In this work, we introduce RayWAM, a World-Action Model that addresses this challenge through Action Raymap, a video-aligned and overcomplete action representation that transforms robot actions into structured RGB maps. By interpreting each end-effector action as the pose of a virtual camera, Action Raymap represents translation, rotation, and gripper states as dense spatial fields, while maintaining a deterministic mapping back to executable robot commands. Based on this formulation, observations and actions are represented in a unified RGB space and encoded using the same frozen pretrained video VAE, projecting robot actions directly into the latent space of pretrained video diffusion models. Unlike prior approaches that adapt video models through additional action-specific modules, RayWAM instead adapts the action representation itself to match the native modality of the pretrained model, enabling more effective exploitation of transferable representations and physical dynamics priors learned from large-scale video data. We evaluate RayWAM on LIBERO, LIBERO-Plus, and VLABench, together with real-world robotic experiments, demonstrating consistent improvements in task success rates, robustness to environmental variations, and real-world generalization across diverse manipulation scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.