Diagnosing and Improving Latent Recovery in Sparse Autoencoders
Abstract
Sparse autoencoders (SAEs) are used to extract interpretable features from neural representations. Their codes assign feature amplitudes used to rank activating examples, attribute model behavior, and set intervention strengths. These applications make amplitude fidelity relevant once feature directions and scales have been fixed. However, reconstruction objectives do not directly measure coefficient accuracy, and interference between co-active features can cause sample-dependent errors that fixed feature-wise rescaling cannot generally remove. Understanding these errors is therefore important for interpreting the codes produced by SAE encoders. We study coefficient recovery in a blind sparse mixing model and characterize amplitude bias in one-pass tied encoders, even with an exact dictionary and correct support. We then compare each code with the code obtained after reconstruction and re-encoding. For tied and untied SAEs, we derive local bounds relating this self-consistency residual to coefficient error under correct and stable support, value-preserving activations, and suitable conditioning. This analysis motivates a self-consistency regularizer that requires one additional encoder evaluation during training and adds no parameters or inference cost. Controlled synthetic experiments show reduced latent recovery error, while CE-bench and SAEBench evaluations provide complementary evidence of improved feature quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.