acceptodds
Under review as a conference paper at ICLR 2027

When Visual Compression Breaks Control: Efficient Multi-View VLA Inference via Retaining Evidence for Control-Aligned Policy Recovery

Abstract

Aggressive visual-token compression is increasingly important for efficient Vision-Language-Action (VLA) inference. However, we find that compression does more than discard visual information: it also induces a representation shift that can substantially misalign a policy trained on dense visual inputs. We therefore formulate efficient multi-view VLA inference as two coupled problems—preserving action-relevant physical evidence and recovering policy behavior under compression-induced representation shift. To this end, we propose RECAP-VLA, short for Retaining Evidence for Control-Aligned Policy Recovery, as a fixed-budget compression-and-recovery framework. RECAP-VLA organizes cross-view tokens into geometry-grounded physical-evidence components and allocates the compact token budget according to action relevance while preserving global context and residual evidence. It further introduces Compression Amplification Profiling (CAP) to localize policy regions most associated with compression-induced control error, and Control-Aligned Counterfactual Policy Adaptation (CAPA) to recover dense-policy behavior using lightweight same-state dense–compact supervision. At a fixed 512-to-256 token operating point, direct compression reduces success from 91% to 46%, whereas RECAP-VLA restores it to 92%. On 200 matched LIBERO Spatial episodes, RECAP-VLA achieves 91.5% success versus 90.5% for the dense policy while reducing median end-to-end latency by 8.44%. The same recipe remains within 2.5 percentage points of dense performance across additional OpenVLA-family checkpoints and LIBERO tasks, with approximately 8.3% median-latency reduction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.