Don't Ignore Them: Scaffold Features Matter for Stable SAE Steering in Diffusion Models
Abstract
Sparse autoencoders (SAEs) enable concept editing in diffusion models by manipulating internal features, and existing methods typically select editing targets based on the target semantics. In contrast, we identify a class of features based on reuse across prompts, without relying on target semantics. Their effects vary with the generated content and are difficult to characterize with a single concept, yet they respond strongly to concept editing. Blocking their responses impairs editing results more than blocking the responses of matched ordinary features. These results show that these features are not high-frequency background unrelated to the target edit, but important internal variables that carry shared generative context. We therefore call them scaffold features. We further incorporate scaffold features into editing and propose Scaffold-Coordinated Feature Editing (SCFE), a method built on existing SAEs that requires no additional training. Using small concept interventions, SCFE estimates spatial responses for the current prompt and adds constrained scaffold-feature updates along the corresponding directions while manipulating the target concept. Experiments show that, under stronger editing, SCFE both increases target-semantic change and improves overall image preservation, and these gains generalize across diffusion-model architectures and their associated SAEs. These findings show that a feature's value for concept editing is not limited to the semantics it represents, as scaffold features selected without relying on target semantics can also improve targeted editing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.