acceptodds
Under review as a conference paper at ICLR 2027

TrajScale-VLA:Trajectory Aligned Scale Autoregressive Action Generation for Vision-Language-Action Models

Abstract

Autoregressive vision-language-action models (VLAs) rely on action tokenization to represent continuous trajectories as discrete sequences for prediction. A key challenge is that fine-grained action tokens preserve control precision but lead to long autoregressive sequences, while hierarchical tokenization can shorten decoding but does not inherently determine how trajectory information should be organized across scales. Existing tokenizers impose an ordering or multi-scale grouping on action tokens, but do not explicitly align each token group with the trajectory structure it is expected to represent. To address this challenge, we propose TrajScale-VLA, which uses trajectory supervision to guide the learning of an action token hierarchy, with early predictions capturing coarse motion and subsequent predictions recovering local detail. We first train a multi-scale residual tokenizer with a multi-scale trajectory refinement (MTR) objective that aligns cumulative reconstructions with targets at progressively finer temporal resolutions. This supervision assigns each scale a distinct reconstruction task within the action hierarchy. We then train a policy to predict the resulting token groups autoregressively across scales and in parallel within each scale, reducing the number of sequential decoding steps. TrajScale-VLA outperforms all compared StarVLA baselines, achieving success rates of 98.4% on LIBERO and 59.8% on RoboCasa-GR1. On real-world USB Pick-and-Insert, which requires coordinated transport and fine alignment, it surpasses the best baseline in each setting by 13.3 and 26.7 percentage points under in-distribution and out-of-distribution conditions. These results highlight the importance of organizing action representations through trajectory supervision for autoregressive robot control.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.