Self-Editing for Safe Image Generation with Unified Multimodal Models
Abstract
Unified multimodal models (UMMs) combine visual understanding with image generation, but this integration does not necessarily translate into more rational generation behavior, e.g., in safety-sensitive scenarios. Specifically, UMMs often fail to prevent safety violations during generation, yet can recognize them in their own outputs. In this paper, we propose **Safety-Guided Self-Editing** (**SGSE**), a framework that connects safety understanding with safe image generation in UMMs. **SGSE** constructs and distills self-correction trajectories using a UMM's own safety assessments. For each image generated by the UMM, **SGSE** uses the same model to identify how safety violations manifest and formulate an editing instruction describing a benign target scene tailored to the image. It then edits toward this target and reassesses the result, repeating the correction if a violation remains. Accepted trajectories supervise both assessment and editing, allowing the model to learn from iterative correction while retaining the same corrective loop at inference time. This process requires neither external safety assessments nor externally provided reference images. Across diverse harmful concepts, **SGSE** reduces the average attack success rate (ASR) on I2P from 31.31% for the base model to 7.84% with at most one edit, achieving safety comparable to that of the untrained loop with up to three edits while requiring 46.0% fewer edits per image. Moreover, **SGSE** generalizes to a concept-erasure task unseen during safety training while largely preserving general image generation capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.