acceptodds
Under review as a conference paper at ICLR 2027

Diversify and Agree: Reliable Surrogate Expansion for Transferable Adversarial Attacks on MLLMs

Abstract

Multimodal large language models (MLLMs) remain vulnerable to transfer-based adversarial attacks, where imperceptible perturbations optimized on accessible surrogate models induce unseen MLLMs to produce attacker-specified responses. While enlarging the surrogate ensemble with diverse representation paradigms improves transferability, the set of accessible surrogates is limited in practice, and collecting or training additional models is costly. In this work, we ask whether richer transferable guidance can be extracted from an existing surrogate set without introducing new models. To this end, we propose DA-Attack, a diversify-and-agree framework comprising two components. Randomized Surrogate Expansion (RSE) diversifies each surrogate by randomly pruning a small fraction of its parameters at every iteration, turning a single model into a stream of online-sampled variants. Since these variants may yield noisy or conflicting gradients, Intra-Surrogate Neighborhood Consensus (ISNC) measures the directional agreement among gradients from the original surrogate and its variants, and assigns greater weight to directions consistently supported across the neighborhood, producing a reliable consensus update. Extensive experiments on eight open-source and four closed-source MLLMs demonstrate that DA-Attack achieves state-of-the-art attack performance. For example, it raises the average targeted attack success rate across the 12 victim models from 70.4% to 82.1% on NIPS 2017, and from 28.5% to 45.6% on Flickr30K, compared with the strongest baseline. The anonymized source code is provided in the supplementary materials.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.