acceptodds
Under review as a conference paper at ICLR 2027

GenShield: Generalizable Visual Prompts for VLM Jailbreak Defense via Dual-View Adversarial Training

Abstract

Vision-Language Models (VLMs) have demonstrated remarkable multimodal capabilities, yet they remain highly vulnerable to jailbreak attacks that bypass their safety alignment and elicit harmful outputs. Defensive visual prompts have recently emerged as a lightweight and plug-and-play defense interface for VLMs. However, existing defensive visual prompts are typically trained against a fixed, hand-curated set of attacks, causing them to overfit to specific attack patterns and scene statistics seen during training, and consequently to transfer poorly to unseen attack strategies or novel visual scenes. We present GenShield, a micro-region adversarial training framework that perturbs only about 2% of the visual input to produce generalizable defensive visual prompts against VLM jailbreak attacks, guided by a dual-view analysis of attack-space coverage (attack view) and prompt information dependency (defense view). From the attack view, we employ beam search over base attack strategies to synthesize the strongest composite attacks in the inner loop, densely covering the attack manifold. From the defense view, we introduce mutual information regularization in the outer loop to strengthen the dependence between the visual prompt and prompt-induced decision changes while reducing their image dependence, promoting consistent defensive effects across visual scenes. Theoretically, we further establish a unified generalization bound for the resulting defensive visual prompts under the dual-view design. Extensive experiments across multiple VLM architectures and attack methods demonstrate that GenShield achieves strong cross-attack and cross-scenario generalization, while largely preserving benign utility and maintaining high visual fidelity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.