Beyond Text-to-Image: Unified Erasure of Sexually Explicit Content in Diffusion Models
Abstract
Diffusion models are increasingly weaponized for sexual abuse, enabling large-scale generation of illicit pornographic content and targeted harassment via fabricated sexualized media depicting real individuals. These threats span four distinct task categories, covering both image and video generation under text-only and image-conditioned inputs. Yet existing defenses and benchmarks are narrowly scoped to text-to-image pipelines, leaving the remaining attack surface exposed. We introduce UniErase-Bench, the first benchmark to systematically evaluate sexually explicit content generation across the full attack surface, and then present UniErase, a unified training-free framework that erases such concepts from the underlying architectures across all task categories. Our key insight is that the feed-forward network (FFN) layers in modern diffusion architectures function as key-value memory units, where sexually explicit concepts are encoded in a subset of channels. Leveraging this property, UniErase first localizes the channels encoding these concepts and then removes them via closed-form weight projection. Extensive experiments on UniErase-Bench demonstrate that UniErase consistently suppresses explicit content generation across all evaluated task categories while preserving the generation quality of benign content, achieving state-of-the-art performance on text-to-image generation and establishing the first unified defense covering the full attack surface. Our code is available at https://anonymous.4open.science/r/UniErase.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.