Exposing Vulnerabilities in Visible-Infrared VLMs: A Unified Adversarial Framework with Cross-Task Transferability
Abstract
Vision–language models (VLMs) have made substantial progress in multimodal understanding and have achieved strong performance across diverse vision–language tasks. However, existing VLM security studies remain largely confined to single-modality settings, predominantly in the visible spectrum, leaving adversarial robustness in visible–infrared (VIS–IR) scenarios underexplored. This omission is consequential for real-world perception systems, where VIS–IR sensing is widely deployed to sustain reliable semantic understanding under challenging imaging conditions. Motivated by this underexplored cross-modal threat setting, we propose a curved-edge fractal adversarial patch (CFPatch), the first unified adversarial patch framework for cross-modal attacks against VIS–IR VLMs. Specifically, CFPatch takes triangular fractal geometry as its base and transforms rigid straight-edged primitives into Bezier-curved elements, preserving fractal self-similarity for multi-scale adversarial effects while enhancing structural expressiveness through smoother contours, richer directional variation, and flexible shape deformation. Complementing this global geometric disruption, we introduce a modality-specific Fraser-spiral rendering mechanism that injects fine-grained texture distortions and misleading perceptual cues tailored to visible and infrared imagery. Through this geometry–appearance coupling, CFPatch forms a dual-level perturbation: the curved fractal geometry disrupts global shape perception, while Fraser-spiral rendering corrupts local appearance interpretation via modality-specific texture interference. We further use expectation over transformation (EOT) to improve patch robustness against common image-level transformations. Extensive experiments show that CFPatch effectively fools VIS–IR VLMs and outperforms standard patch baselines in both attack effectiveness and robustness. Moreover, adversarial samples optimized for the zero-shot classification task transfer well to image captioning and visual question answering (VQA), demonstrating the strong cross-task transferability and generalizability of CFPatch across downstream tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.