acceptodds
Under review as a conference paper at ICLR 2027

Discrete Text-Conditioned Universal Multimodal Attacks on Vision–Language Models

Abstract

Vision-language pre-trained (VLP) models support diverse multimodal tasks but remain vulnerable to universal adversarial perturbations. Continuous textual surrogates can create a mismatch between optimization and the discrete edits applied at inference, while a shared word has context-dependent effects. We propose an optimization-based framework that combines online vocabulary-constrained text learning with edited-caption conditioning: the current shared word is inserted into complete captions, which are re-encoded to condition universal image-perturbation learning. Representation-level disruption and bidirectional cross-modal competition target image–text correspondence, with auxiliary fusion guidance when supported by the source architecture. Experiments span six VLPs configurations and three vision-language tasks. Extensive experiments demonstrate competitive white-box attack performance, and effective transfer to downstream vision-language tasks without further optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.