ActionCodec2: A Streaming Action Tokenizer
Abstract
In autoregressive vision-language-action (AR-VLA) models, action tokenizers map continuous actions to discrete tokens and back to executable actions. Most existing tokenizers are chunkwise, allowing the robot to react only at chunk boundaries. AR models can instead support streaming tokenization, where each action is decoded as soon as its tokens are available. Existing streaming designs, however, either require many tokens or lack distribution-independent reconstruction guarantees, while small errors in incremental actions can accumulate into trajectory drift. We therefore seek the minimum token rate subject to bounded reconstruction error and a user-specified trajectory tolerance. Under the group structure of the action space, this becomes a covering problem. We prove fixed motion primitives optimal among one-label-per-step decoders, with an exact closed-form solution for translation and constant-factor-tight bounds for orientation. We further compress the primitive stream losslessly with Set-BPE, a byte-pair encoding over unordered sets. The resulting ActionCodec2 requires no gradient training. Fitted on k hours from embodiments, it encodes one second of motion in – tokens on average. Using the same VLA backbone, ActionCodec2 outperforms the compared tokenizers on LIBERO (), LIBERO-Plus (), and SimplerEnv (), and reaches task progress across four zero-shot Franka tasks. More results are available at https://actioncodec2-anon.pages.dev.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.