XSPA: Crafting Imperceptible X-Shaped Sparse Adversarial Perturbations for Transferable Attacks on VLMs
Abstract
Vision-language models (VLMs) rely on a shared visualtextual representation space to support tasks such as zeroshot classification, image captioning, and visual question answering (VQA). Although this shared space enables strong cross-task generalization, it also creates a critical vulnerability: subtle visual perturbations may propagate through the common embedding space and induce correlated semantic failures across tasks. This issue is especially concerning in interactive and decision-support scenarios, yet it remains unclear whether VLMs are still fragile under highly constrained, sparse, and geometrically fixed perturbations. We propose Xshaped Sparse Pixel Attack (XSPA), an imperceptible structured attack that restricts perturbations to two intersecting diagonal lines. Compared with dense perturbations and flexible localized patches, XSPA operates under a much stricter attack budget, offering a stringent test of VLM robustness. Within this sparse support, XSPA jointly optimizes a classificationoriented objective, cross-task semantic guidance, and regularization on perturbation magnitude and linewise smoothness, inducing both transferable misclassification and semantic drift in captioning and VQA while preserving visual subtlety. Under the default setting, XSPA modifies only about 1.04% of image pixels. Experiments on the COCO dataset show that XSPA consistently degrades performance across zero-shot classification, image captioning, and VQA. Zero-shot accuracy drops by 52.33 points on OpenAI CLIP ViT-L/14 and 67.00 points on OpenCLIP ViT-B/16, while GPT-4-evaluated caption consistency decreases by up to 58.60 points and VQA correctness by up to 44.25 points. These results show that even highly sparse and visually subtle perturbations with fixed geometric priors can substantially disrupt cross-task semantics in VLMs, revealing an important robustness gap in current multimodal systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.