GenShield: Generalizable Visual Prompts for VLM Jailbreak Defense via Dual-View Adversarial Training
Abstract
Vision-Language Models (VLMs) have demonstrated remarkable multimodal capabilities, yet they remain highly vulnerable to jailbreak attacks that bypass their safety alignment and elicit harmful outputs. Defensive visual prompts have recently emerged as a lightweight and plug-and-play defense interface for VLMs. However, existing defensive visual prompts are typically trained against a fixed, hand-curated set of attacks, causing them to overfit to specific attack patterns and scene statistics seen during training, and consequently to transfer poorly to unseen attack strategies or novel visual scenes. We present GenShield, a micro-region adversarial training framework that perturbs only about 2% of the visual input to produce generalizable defensive visual prompts against VLM jailbreak attacks, guided by a dual-view analysis of attack-space coverage (attack view) and prompt information dependency (defense view). From the attack view, we employ beam search over base attack strategies to synthesize the strongest composite attacks in the inner loop, densely covering the attack manifold. From the defense view, we introduce mutual information regularization in the outer loop to strengthen the dependence between the visual prompt and prompt-induced decision changes while reducing their image dependence, promoting consistent defensive effects across visual scenes. Theoretically, we further establish a unified generalization bound for the resulting defensive visual prompts under the dual-view design. Extensive experiments across multiple VLM architectures and attack methods demonstrate that GenShield achieves strong cross-attack and cross-scenario generalization, while largely preserving benign utility and maintaining high visual fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.