acceptodds
Under review as a conference paper at ICLR 2027

Spatially Anchored Queries for Efficient Vision-Language-Action Models

Abstract

Vision-language-action (VLA) models typically process hundreds of visual tokens at every control step, making visual token compression important for efficient robot control. However, aggressive compression is not only about reducing token count, but also about how limited representational capacity is allocated across the scene. Existing pruning methods retain discrete visual tokens, while continuous sampling methods represent each compact token by a single spatial coordinate, potentially discarding broader spatial evidence needed for action. We introduce Anchored Query Heatmaps (AQH), a structured visual bottleneck that represents each compact token with a persistent spatial anchor and a task-conditioned spatial posterior. The posterior guides a precise continuous read from the dense visual feature map while retaining broader spatial context. To capture this context efficiently, we further introduce a Recursive Conditional Feature Field (RCFF) that performs posterior-aware aggregation in a shared low-rank spatial field. With only 32 of the original 512 visual tokens, AQH achieves 95.65% average success on LIBERO-2000 and outperforms existing compression methods under the same token budget. A controlled 32-token configuration yields a 1.92 end-to-end policy speedup over the dense reference. AQH also shows consistent gains on LIBERO-Plus and RoboTwin, demonstrating that spatially structured representations enable aggressive visual compression while preserving the evidence needed for action.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.