ActWeave: Efficient Action Generation for Vision-Language-Action Models via Hierarchical Autoregression and Wavefront Decoding
Abstract
Autoregressive vision-language-action (VLA) models have demonstrated strong robotic manipulation capabilities. However, fine-grained action tokenization can produce long sequences, making token-wise decoding through large vision-language models (VLMs) expensive. We introduce **ActWeave**, a hierarchical autoregressive framework that shifts VLM computation from tokens to structured action patches. A deterministic discrete cosine transform (DCT) tokenizer retains all quantized coefficients in an action-dimension–frequency grid without a learned codebook. The VLM-based Thinker models inter-patch dependencies, while a lightweight Weaver performs **wavefront decoding**, predicting each anti-diagonal in parallel conditioned on earlier wavefronts. Direct VLM fine-tuning without robot-data pretraining yields competitive LIBERO and SimplerEnv performance. Controlled LIBERO comparisons show a 17.9× full-chunk speedup over flat global autoregression. Wavefront decoding reduces latency by 36.0% relative to raster local decoding at comparable success rates. VLABench reconstruction-and-replay experiments show that the tokenizer preserves control-relevant information under the evaluated distribution shift. These results demonstrate efficient autoregressive action generation without shortening the action-token sequence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.