acceptodds
Under review as a conference paper at ICLR 2027

Flexible Action-Oriented Image Tokenization

Abstract

Vision-language-action (VLA) models commonly process hundreds of patch-level visual features at every control step, creating long visual contexts. These visual contexts constitute the largest computational bottleneck in VLAs. Existing VLA compression methods may discard control-relevant information at low token budgets or learn representations tied to a predetermined token count, limiting inference-time adaptation. We introduce Flexi-Act, a visual tokenizer that compresses patch features into a compact, ordered sequence of continuous tokens. This sequence preserves control-relevant information under aggressive compression and supports different visual-token budgets through its prefixes. To learn this ordering, we pretrain the tokenizer with nested dropout and a temporary reconstruction decoder that reconstructs frozen visual-encoder features from sampled prefixes, encouraging earlier tokens to capture broadly useful information while later tokens specialize. We then integrate the pretrained tokenizer into a downstream VLA and jointly fine-tune the tokenizer and policy across prefix lengths, allowing the same policy to switch between supported budgets at inference without retraining. At 16 tokens per camera view, Flexi-Act maintains strong control performance across LIBERO, LIBERO-Plus, and SIMPLER while using 16× fewer visual tokens and 4.5×–5.8× lower estimated FLOPs than OpenVLA-OFT. On LIBERO and LIBERO-Plus, it maintains near-lossless performance using only 1 visual token per camera view.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.