acceptodds
Under review as a conference paper at ICLR 2027

Perception over Action? Rethinking Adversarial Robustness in Vision-Language-Action Models

Abstract

Visual adversarial patches can disrupt perception in vision–language–action (VLA) models and prevent manipulation tasks from being completed. We systematically compare robustness adaptation strategies to examine how their training objectives and adaptation scope relate to unseen-attack robustness and clean-task preservation. Specifically, we compare action-supervised adaptation of the entire policy, termed VLA Tune, with Visual Encoder Tune, which aligns visual features with a clean reference while retaining the original downstream policy. Across four LIBERO suites, Visual Encoder Tune achieves 66.6% mean clean success, compared with 55.0% for VLA Tune and 63.5% for the original policy. On LIBERO-Spatial, its mean success under three unseen attacks reaches 60.0%, versus 47.2% for VLA Tune and 0.0% for a visual-augmentation baseline. Representation diagnostics show closer recovery of clean-reference features under attack and less drift on clean observations. The advantage in task success also appears under transient and persistent patch exposure, although attacks optimized against the final defender substantially reduce performance. These findings support visual-encoder adaptation with clean-reference alignment as an effective strategy for improving transfer robustness while preserving acquired manipulation skills.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.