acceptodds
Under review as a conference paper at ICLR 2027

Safe Latent Control Flow: Safe Diffusion Sampling via Chance-Constrained Control

Abstract

Pretrained text-to-image diffusion models can generate unsafe content, including nudity, violence, and memorized training images. Existing training-free safety methods steer the sampling trajectory away from unsafe regions, but they typically use a guidance window fixed in advance, and none of them states how often the outputs it returns remain unsafe. We propose Safe Latent Control Flow (SLCF), a training-free sampler that states the safety requirement as a probability and switches the control off on each sample once that requirement is met. We formulate safe generation as a chance-constrained stochastic optimal control problem that limits the probability of unsafe outputs while staying close to the pretrained sampling process. We derive the control from the denoiser's existing prediction at each step, and the same prediction decides when the sample would end up safe without control, against a threshold calibrated once per task. We prove that the samples it switches off meet the probability constraint under assumptions on the calibration. On three safe generation benchmarks, SLCF achieves the lowest unsafe rate among training-free baselines and preserves benign generations better, and the bound we certify holds under an external evaluator on every prompt set.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.