Spatially Anchored Queries for Efficient Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models typically process hundreds of visual tokens at every control step, making visual token compression important for efficient robot control. However, aggressive compression is not only about reducing token count, but also about how limited representational capacity is allocated across the scene. Existing pruning methods retain discrete visual tokens, while continuous sampling methods represent each compact token by a single spatial coordinate, potentially discarding broader spatial evidence needed for action. We introduce Anchored Query Heatmaps (AQH), a structured visual bottleneck that represents each compact token with a persistent spatial anchor and a task-conditioned spatial posterior. The posterior guides a precise continuous read from the dense visual feature map while retaining broader spatial context. To capture this context efficiently, we further introduce a Recursive Conditional Feature Field (RCFF) that performs posterior-aware aggregation in a shared low-rank spatial field. With only 32 of the original 512 visual tokens, AQH achieves 95.65% average success on LIBERO-2000 and outperforms existing compression methods under the same token budget. A controlled 32-token configuration yields a 1.92 end-to-end policy speedup over the dense reference. AQH also shows consistent gains on LIBERO-Plus and RoboTwin, demonstrating that spatially structured representations enable aggressive visual compression while preserving the evidence needed for action.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.