acceptodds
Under review as a conference paper at ICLR 2027

SharedSAE: A Unified Sparse Dictionary across U-Net Blocks and Denoising Time for Diffusion Models

Abstract

Concept representations in diffusion models drift across network depth and denoising time, yet existing sparse autoencoders offer only per-block or per-timestep dictionaries, leaving no single feature dictionary for localization and intervention in the network. We present Shared Sparse Autoencoder (SharedSAE), which learns one sparse dictionary shared across four U-Net attention blocks and all 50 denoising steps of SDXL 1.0. A block-specific low-rank input adapter aligns the residual distribution of each block before the shared encoder, and a temporal conditioning module with a shared timestep branch and a block-timestep correction branch absorbs temporal non-stationarity. Despite spanning both axes, SharedSAE preserves a coherent dictionary structure. Built on the shared dictionary, localized feature suppression reduces the number of NudeNet-positive generations in the I2P detector-positive subset from 139 to 1. Tracing sparse features across the trajectory further reveals that concept representations are redistributed across U-Net blocks during generation, shifting from the earliest downsampling block toward intermediate and later blocks as denoising proceeds. SharedSAE thus provides a unified, interpretable, and intervention-ready representation space for diffusion models, enabling concept analysis and erasure at any covered block and denoising step.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.