Detecting and Stopping NSFW Generation via Denoising Dynamics
Abstract
Text-to-image diffusion models can synthesize not-safe-for-work (NSFW) content. Existing detection methods primarily assess the content of prompts, completed images, or intermediate states. However, prompt filters are vulnerable to obfuscated intent, while post-generation classifiers incur the full generation cost before detection. In contrast, in-generation detectors inspect intermediate features but rely primarily on static cues from the semantic content of individual states, leaving activation dynamics across denoising steps underexplored. We therefore model harmful dynamics, i.e., temporal patterns of activation motion associated with unsafe content, as a safety signal. Based on this signal, we formulate detection as a constrained sequential decision problem, aiming to stop unsafe generations early while balancing detection errors and sampling cost. Our approach follows an “observe, summarize, and stop” workflow: it observes activation motion during denoising, summarizes its history into risk scores, and stops generation based on accumulated risk. We conduct comprehensive experiments across four diffusion backbones and three benchmarks. Our approach significantly improves in-domain and cross-domain detection, with relative AUC gains of up to 14.7% over the strongest baseline, demonstrating the value of denoising dynamics as a safety signal. In addition, dynamics-based monitoring generalizes to unseen prompt distributions across different diffusion architectures and enables early intervention in unsafe generation with minimal overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.