acceptodds
Under review as a conference paper at ICLR 2027

Granularity-aware Semantic Transfer Learning for Targeted MLLM Attacks

Abstract

Transfer-based targeted attacks provide an important means of evaluating the robustness of closed-source multimodal large language models (MLLMs). Recent advances have improved transferability by exploiting increasingly fine-grained and hierarchical visual representations. However, existing objectives are still largely designed around the representations being aligned. Local views, patch representations, and intermediate features are often optimized through separate objectives, without explicitly connecting them through shared target semantics. As a result, semantically related signals can become fragmented across visual granularities. To address this issue, we introduce **Granularity-Aware Semantic Transfer (GAST)**, which leverages target semantics to coordinate surrogate supervision across different visual granularities. Specifically, Semantic Scope Construction (SSC) establishes distinct spatial scopes for different semantics in the target image, providing localized guidance for subsequent alignment. Correspondence-Aware Full-Patch Alignment (CFPA) then preserves the complete set of patch representations within each semantic scope and establishes optimal one-to-one correspondences between adversarial and target patches. Furthermore, Hierarchy-Wide Semantic Alignment (HWSA) propagates the aligned semantic signals across the intermediate representation hierarchy, enabling representations at different depths to jointly contribute to the transfer objective. Black-box evaluations on ten open-source and closed-source MLLMs demonstrate that GAST outperforms the second-best method by 12%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.