Revealing Hidden Intent for Defending Text-to-Image Models against Adversarial Attacks
Abstract
Adversarial attacks against Text-to-Image (T2I) models can induce unsafe imagery while concealing underlying harmful intent through obfuscated perturbations. Existing defenses primarily mitigate such threats either by suppressing predefined unsafe patterns via interventions in the T2I generation process or by evaluating manipulated prompts with auxiliary signals, such as latent concept similarity and reconstruction discrepancy. However, their effectiveness can deteriorate when harmful content falls outside the targeted space or when such cues provide insufficient evidence of semantics rendered opaque by adversarial perturbations, underscoring the need to reveal the semantics these manipulations obscure. In this paper, we introduce EEGIS, a T2I defense framework that surfaces the latent intent embedded in adversarial prompts as semantically coherent descriptions for reliable and interpretable safety assessment. EEGIS reconstructs the concealed semantics of lexically corrupted expressions by inferring the visual content evoked by each perturbed token from the diffusion noise during early denoising. Building on the recovered semantics, it further grounds implicit implications into visually consistent explicit alternatives, exposing harmful intent that remains concealed beyond surface-level lexical distortions. Extensive experiments on NSFW-200 and T2I-RiskyPrompt across diverse attack methods demonstrate that EEGIS effectively captures adversarial intent while maintaining low false positive rates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.