acceptodds
Under review as a conference paper at ICLR 2027

Deep Localized Concept Erasure via Visually-Cued Bregman Preference Optimization

Abstract

Concept erasure in text-to-image diffusion models demands both spatial locality to preserve benign contexts and deep removal to withstand adversarial reactivation. Existing methods fall short on the two dimensions: text-centric approaches leave undesirable visual features intact, while image-centric preference learning (e.g., DPO) suffers from global entanglement and gradient vanishing, resulting in superficial marginal ranking rather than persistent erasure of undesirable concepts. To overcome these bottlenecks, we propose a Visually-Cued Bregman Preference Optimization framework for deep localized concept erasure. First, we enforce strict spatial locality by grounding preference optimization on masked, minimally-edited safe-unsafe image triplets. Second, to achieve deep erasure, we deliberately expose visual cues of undesirable concepts to the model during fine-tuning, forcing the model to thoroughly erase the concept under strong adversarial visual stimuli. Finally, we further propose to use Bregman Preference Optimization (BPO) for concept erasure and generalize it to diffusion models. We analytically demonstrate that Diffusion BPO provides non-vanishing optimization signals, sustaining continuous concept suppression well beyond the limitations of DPO. Extensive experiments demonstrate that our framework achieves superior erasure efficacy, robust resistance against optimized textual and multimodal reactivation attacks, and high-fidelity locality preservation compared to representative baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.