Sparse Autoencoders Reveal Temporally Stable Medical Semantics in Diffusion Models
Abstract
Sparse Autoencoders (SAEs) are useful for discovering sparse directions in generative-model activations, but diffusion activations vary strongly across denoising timesteps. An SAE trained on raw multi-timestep activations can therefore mix class-relevant medical structure with noise-stage effects, making the same sparse direction unstable over time. We introduce Piecewise DoD-SAE, a simple training recipe that reconstructs scale-normalized local Direction-of-Deviation (DoD) residuals between neighboring latent trajectory states rather than raw hidden states. The objective is motivated by residual-scale diagnostics along the denoising path and is designed to expose sparse directions whose medical responses persist across timesteps. We evaluate temporal stability with top- latent IoU, image-space steering consistency, and matched classifier scoring of guided images. Across three liver MRI benchmarks, Piecewise DoD-SAE improves cross-timestep stability and steering consistency over SAE baselines under matched protocols. The results support a calibrated claim: piecewise DoD helps SAEs reveal temporally stable medical latent responses in diffusion models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.