FlowSAE: Non-Invasive Sparse Autoencoding with Lifted Invertible Flows
Abstract
Sparse autoencoders (SAEs) are used both to interpret neural network representations and to intervene on them, but a standard SAE's decoder is only an approximate inverse of its encoder, so reconstruction error contaminates any causal effect an intervention is meant to isolate. Prior work compensates for this after the fact, by carrying the error as an external residual channel or by normalising away its damage, rather than removing it. We introduce FlowSAE, to our knowledge the first nonlinear, multi-layer flow-based SAE whose decoder is the exact inverse of its encoder. FlowSAE lifts the activation into a larger space, applies an invertible flow, sparsifies, and inverts. Top- sparsification is the only lossy step. When no feature is removed, the input is recovered exactly for any learned parameters, not only after training. On ImageNet ViT features this holds to numerical precision. It also removes a problem we find in classical SAEs, where uncorrected single-feature ablations can invert the true attribution ranking. At the deployed sparsity, however, exactness does not carry over. The nonlinear flow reconstructs better than a linear exact control, but its feature effects become non-additive, so the standard error-term correction, which is exact for any linear decoder, no longer recovers them. Exact reconstruction adds no asymptotic cost per forward pass, only a constant factor. Making sparse interventions trustworthy is a separate problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.