Rethinking Transferable Targeted Attacks on Multimodal Large Language Models via Causal Intervention
Abstract
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding, yet their robustness against targeted adversarial manipulation remains largely unexplored. Existing approaches on transferable targeted attacks mainly optimize target-related feature representations but overlook which visual components truly drive the transfer toward the target semantics. To address this limitation, we introduce a causal-guided framework that leverages target-aware counterfactual intervention to uncover the visual factors contributing to adversarial transfer. To achieve efficient estimation, we develop a gradient-based approximation strategy that evaluates the causal influence of visual components. The identified target-relevant components are then incorporated into a causal-guided feature optimization objective to improve alignment with the desired semantic representation. Extensive experiments across various MLLMs demonstrate that our approach consistently enhances targeted transferability over existing attack methods, validating the effectiveness of causal guidance for generating transferable adversarial examples.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.