Rethinking Adversarial Patch Attacks on VLA: From Pixel Space to Physical Reality
Abstract
Vision-Language-Action (VLA) models are increasingly deployed in safety-critical manipulation tasks, making adversarial robustness a security-critical concern. Recent studies report that adversarial patches can reduce task success rate to near zero. But how much of this attack strength survives when a patch is fixed to a physical surface? We revisit this question by separating patch generation from patch evaluation. Pixel-space compositing places a patch directly in the input image, whereas a fixed physical texture must produce all of its appearances through camera projection. We cross these two settings in a study on OpenVLA, OpenVLA-OFT, and . Switching from pixel-space to geometry-consistent evaluation raises task success rate in all 24 combinations, by 3.0-69.2 percentage points. Re-optimizing under geometric constraints then recovers attack strength in 23 of these combinations. For example, OpenVLA's success rate on LIBERO-Spatial under UADA attack changes from 0.2% to 52.0% when evaluation changes, and to 17.4% after re-optimization. Printed patches also disrupt a physical robot task. Finally, we examine how the same constraints affect adversarial training, finding substantial gains on VLA models. Together, these results show why physical patch risk should be evaluated under an explicit deployment model, with attack generation and evaluation considered separately.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.