Training-Free Concept Erasure via Text-Projection Weight Editing for Visual Autoregressive Models
Abstract
Visual autoregressive (VAR) models have emerged as a competitive paradigm for high-resolution text-to-image generation. However, their growing generative capability raises concerns about undesirable content. Concept erasure aims to remove such concepts from pretrained models while preserving their general generation capability. Existing approaches to concept erasure in VARs either require time-consuming fine-tuning or introduce additional overhead through inference-time intervention. To address these limitations, we propose **TRACE** (**TRA**ining-Free **C**oncept **E**rasure), a training-free framework based on text-projection weight editing for VARs. Specifically, TRACE edits the linear layer of the text-projection module and traces concept-induced responses across different scales to construct scale-specific concept subspaces. TRACE then applies closed-form low-rank corrections to the text-projection weights to modulate components of the target concept from the projected text representations along these subspaces. We further introduce TRACE-Lite, a lightweight variant, which uses a single concept subspace and weight edit across all scales, completing concept erasure in about 3.40 seconds on a single GPU. Extensive experiments demonstrate that TRACE effectively erases target concepts while preserving generation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.