Visual Exclusivity Attacks: Automatic Multimodal Red Teaming via Agentic Planning
Abstract
Many multimodal red teaming methods use images as wrappers for malicious payloads via typography or adversarial noise. A complementary threat arises when the image supplies information needed for an image-specific harmful answer. We formalize Visual Exclusivity (VE), an Image-as-Basis criterion defined relative to specified natural-language descriptions and budgets. To study this threat, we propose Multimodal Multi-turn Agentic Planning (MM-Plan), a framework that reframes jailbreaking from turn-by-turn reaction to global plan synthesis. MM-Plan trains an attacker planner to synthesize comprehensive multi-turn strategies, optimized via Group Relative Policy Optimization (GRPO), learning text queries and the use of predefined visual operations without human-annotated attack trajectories. We introduce VE-Safety, a dataset of 440 human-curated instances across 15 safety categories, screened for dependence on visual content such as technical schematics. Across eight evaluated MLLMs, MM-Plan outperforms the compared baselines. It achieves 46.3% attack success against Claude 4.5 Sonnet and 13.8% against GPT-5. These findings reveal that frontier models remain vulnerable to learned multimodal interaction on tasks curated for visual dependence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.