acceptodds
Under review as a conference paper at ICLR 2027

Being Trapped then Mitigated: Defending Malicious Jailbreak Prompts via Embedded Honeypots in Text-to-Image Models

Abstract

The rapid popularity of Text-to-Image (T2I) models has raised widespread concerns about their non-ethical application for generating inappropriate or harmful content. To prevent their misuses, several methods for safety alignment, also known as concept erasure, have been developed, yet vulnerable to adversarial attacks, which exploit optimized jailbreak prompts to break the safety barriers, regenerating undesired content. In this paper, distinct from previous arts that patch model vulnerabilities via adversarial training and external defensive tools, we introduce a proactive defense framework, HoneyT2I, which exploits embedded honeypots as artificial vulnerabilities to manipulate jailbreak optimization, making the jailbreak behavior of resulting prompts to be detected and mitigated easier. Specifically, the honeypots of unsafe concepts are embedded into pretrained T2I models with the backdoor learning paradigm, ensuring target concepts are regenerated once the corresponding trigger emerges. Moreover, Singular Vector Decomposition is utilized to extract honeypot and difference signatures from prompt collections containing target concepts and their backdoor versions, which further detect and mitigate jailbreak behavior at the inference time. Extensive experiments show that HoneyT2I effectively defend jailbreak prompts generated by various attacks without sacrificing generation quality, while possessing the promising scalability, i.e., being integrated with existing alignment methods and applied to advanced T2I models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.