acceptodds
Under review as a conference paper at ICLR 2027

Detecting and Stopping NSFW Generation via Denoising Dynamics

Abstract

Text-to-image diffusion models can synthesize not-safe-for-work (NSFW) content. Existing detection methods primarily assess the content of prompts, completed images, or intermediate states. However, prompt filters are vulnerable to obfuscated intent, while post-generation classifiers incur the full generation cost before detection. In contrast, in-generation detectors inspect intermediate features but rely primarily on static cues from the semantic content of individual states, leaving activation dynamics across denoising steps underexplored. We therefore model harmful dynamics, i.e., temporal patterns of activation motion associated with unsafe content, as a safety signal. Based on this signal, we formulate detection as a constrained sequential decision problem, aiming to stop unsafe generations early while balancing detection errors and sampling cost. Our approach follows an “observe, summarize, and stop” workflow: it observes activation motion during denoising, summarizes its history into risk scores, and stops generation based on accumulated risk. We conduct comprehensive experiments across four diffusion backbones and three benchmarks. Our approach significantly improves in-domain and cross-domain detection, with relative AUC gains of up to 14.7% over the strongest baseline, demonstrating the value of denoising dynamics as a safety signal. In addition, dynamics-based monitoring generalizes to unseen prompt distributions across different diffusion architectures and enables early intervention in unsafe generation with minimal overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.