acceptodds
Under review as a conference paper at ICLR 2027

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

Abstract

World Action Models (WAMs) leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control. Existing numerical actions fail to satisfy the former, and prior visual action representations overlook the temporal motion structure across frames. We address this issue with FlowWAM, a dual-stream diffusion framework that adopts optical flow as a unified, video-native action representation. Rather than attaching actions to the generator through a separate action space, FlowWAM natively internalizes them as a parallel video stream. Since flow videos share the same format as RGB videos and encode rich per-pixel displacement, FlowWAM jointly models both within a shared pretrained video generator, naturally implementing two modes of WAMs. In policy mode, FlowWAM generates flow for action prediction, while in world-model mode, it uses target flow sequences to guide future video generation. Moreover, since learning flow requires no action labels, FlowWAM can leverage large-scale action-unlabeled videos for pretraining. In policy mode, FlowWAM outperforms both VLA and WAM baselines on RoboTwin (93.7% average success), with action-unlabeled video pretraining notably boosting generalization to unseen randomized scenes, and surpasses dynamic-manipulation specialists on DOMINO (41.3% success). In world-model mode, its dense motion conditioning achieves the highest overall score on WorldArena, bringing an 18.4% relative improvement in trajectory accuracy. More results are available at https://anonymous-flowwam.github.io.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.