Spot the Difference: Efficient VLA Inference via Differential Coding
Abstract
Vision-language-action (VLA) models inherit the perception and reasoning of large vision-language models (VLMs), but they pay for it on every policy query: a multi-billion-parameter VLM re-encodes each new observation, even though consecutive frames of a manipulation rollout are highly redundant. We propose Differential VLA, which applies a principle shared by differential coding in communications, predictive coding in cognitive science, and I-frame/P-frame video compression: *encode a reference once, and just spot the difference afterwards*. Differential VLA runs the VLM (System 2) on a reference frame only once every control steps and caches its tokens. In between, a lightweight *differential encoder* lets the features of the current frame attend to those of the reference frame and compresses the residual into a handful of *difference tokens*, which are appended to the cached tokens that condition the diffusion-transformer action head (System 1). Built on GR00T N1.7 with , Differential VLA achieves an average success rate of 93.5% on LIBERO, which is within 1.0 points of the baseline, while calling the VLM 8× less often. An amortized query costs 1.8× fewer FLOPs and runs 1.27× faster than an equally optimized baseline. On SimplerEnv, Differential VLA stays within 6.6 points of the baseline while calling the VLM 16× less often, and on two real-world tasks with an SO-101 arm, it matches the baseline's success rate with 1.22 to 1.27× lower latency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.