Enhancing Transferable Targeted Attacks on Closed-Source MLLMs with Progressive Semantic Chain
Abstract
Transfer-based targeted attacks on commercial black-box multimodal large language models (MLLMs) remain challenging because attackers can access only the model’s inputs and outputs. Existing methods primarily rely on visual semantic alignment in the embedding space to guide adversarial example generation, but the resulting perturbations often lack sufficiently explicit semantic cues, which limits their transferability. In contrast, textual descriptions can provide more explicit and fine-grained semantic cues for adversarial optimization. Motivated by this insight, we propose Progressive Semantic Chain Attack (PSC-Attack), which introduces a semantic chain to provide explicit semantic guidance for adversarial example generation. Specifically, PSC-Attack decomposes the target semantics into a sequence of concise semantic nodes that capture explicit semantic cues and progressively guide the optimization in the shared vision-language embedding space. Furthermore, we introduce a Minimum-Redundancy Semantic Chain to adaptively select more explicit and fine-grained target semantic nodes while suppressing redundant guidance, thereby enhancing attack transferability. Extensive experiments on closed-source commercial MLLMs demonstrate that PSC-Attack consistently outperforms existing targeted transferable attack methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.