acceptodds
Under review as a conference paper at ICLR 2027

ASSIST-VLA: ACCELERATING VISION-LANGUAGE- ACTION INFERENCE VIA MULTI-LEVEL REDUN- DANCY REDUCTION

Abstract

Vision-language-action (VLA) models enable robots to translate visual observa- tions and language instructions into actions, but repeated multimodal inference in- curs substantial latency during closed-loop control. We introduce ASSIST-VLA, a training-free framework that accelerates VLA inference by reducing redundancy at three levels: control cycles, visual inputs, and model structure. At the control- cycle level, consistency across successive action plans guides the reuse of cached cognition representations, allowing selected steps to bypass visual encoding and language-backbone inference. At the visual-input level, instruction-conditioned semantic relevance and local spatiotemporal cues guide token pruning to retain task-relevant information. At the model-structure level, we draw on the J-space perspective, which analyzes internal representations through their influence on model outputs, and extend it to VLA action outputs. By combining Jacobian mappings with channel activation statistics, we assess the influence of MLP chan- nels on predicted actions and structurally prune low-influence channels. Together, these components reduce the frequency and cost of full multimodal inference. On four manipulation tasks in SimplerEnv with CogACT-Base, ASSIST-VLA reduces average end-to-end inference latency from 229.1 ms to 158.5 ms per control step, achieving a 1.45× speedup with an average task success rate of 67.8%, compared with 67.1% for the original model. These results show that reducing redundancy across multiple levels can accelerate closed-loop VLA inference while maintain- ing task performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.