acceptodds
Under review as a conference paper at ICLR 2027

TIGA: Improving Targeted Transferability on MLLMs through Global-Regional Alignment and Neighborhood Gradient Fusion

Abstract

Black-box targeted adversarial transfer attacks on Multimodal Large Language Models (MLLMs) aim to induce outputs of attacker-specified semantics through carefully crafted adversarial examples without accessing any internal information of the victim models. Existing methods typically generate adversarial examples by aligning adversarial and target features on accessible surrogate models, utilizing model ensembles and random cropping to further enhance adversarial transferability. However, feature alignment of previous work may overlook coarse regional information relevant to language generation, while gradients evaluated at individual inputs may provide insufficient optimization guidance under local input variations. To address these limitations, we propose the TIGA framework, which improves targeted transferability from both feature and optimization perspectives. Specifically, Global-Regional Collaborative Alignment combines global features with patch representations within each region of a grid, preserving spatial information without strict patch-wise correspondence. Moreover, Neighborhood-Aware Gradient Fusion combines normalized gradients from the current adversarial input and a noise-perturbed neighbor, providing complementary guidance for targeted alignment under local input variations. Experiments on six closed-source and six open-source MLLMs demonstrate that TIGA outperforms existing methods in both attack success rates and output semantic similarities, with particularly pronounced gains on the closed-source models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.