acceptodds
Under review as a conference paper at ICLR 2027

Ordered Action Transport for Generative Robot Policies

Abstract

Generative robot policies, including vision-language-action models (VLAs) and world-action models (WAMs), generate an action chunk by jointly refining a sequence of action tokens, one per future action. These tokens inherit rotary position encoding (RoPE), which relates two tokens by how many steps apart they are. However, existing positional encodings represent action-token relations primarily by their positions in the sequence, overlooking the motion induced by the actions are taken and in what order. We propose Action-Transport RoPE (AT-RoPE), which advances the encoding at each step by a rotation determined by the action rather than by a fixed one. Each current noisy action parameterizes an orthogonal operator, and these operators are composed in execution order, so that the prefix transformations of two tokens differ by the composition of the actions between them. AT-RoPE reduces to RoPE-style rotation when the step is action-independent and to an order-insensitive summed prefix when the operators commute. For backbones with rotary attention, a complementary pathway further conditions query-key matching on the actions being compared. Across four VLA and WAM backbones, AT-RoPE improves success on LIBERO and RoboTwin 2.0 for every backbone, with consistent gains on long-horizon tasks, and improves zero-shot robustness on LIBERO-Pro and LIBERO-Plus. Ablations confirm that ordered composition outperforms prefix-free and summed-prefix alternatives, and rollouts with AT-RoPE are smoother.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.