Adaptive Inpainting Masking from Unconditional Concept Erasure Bias for Safe Inpainting
Abstract
Recent diffusion models enable high-fidelity image generation, editing, and inpainting, but they also reopen a subtle attack surface: an adversary can specify a narrowly localized inpainting mask so that the target concept is underrepresented in the masked region, thereby bypassing Anti-editing Concept Erasure (ACE). This mismatch stems from a design gap—current erasure methods operate on the denoising pathway but do not verify whether the user-specified inpainting mask actually covers the concept to be removed. We address this gap by introducing Adaptive Inpainting Masking (AIM) from Unconditional Concept Erasure Bias, a training-free pipeline that turns the ACE-imposed unconditional bias into a pre-inpainting mask refinement signal. We first isolate the object into an OOD-like view, then run deterministic unconditional DDIM inversion through the ACE model, and measure reconstruction consistency. Regions whose visual evidence aligns with the ACE-style unconditional bias are interpreted as containing the erased concept and are automatically merged into the inpainting mask. The refined mask thus captures concept-bearing areas that an attacker attempted to exclude, enabling the downstream inpainting step to remain concept-safe without any explicit text labels. Our approach effectively extends diffusion-based concept erasure to the inpainting setting, closing a practical evasion channel while preserving the training-free nature of existing ACE frameworks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.