acceptodds
Under review as a conference paper at ICLR 2027

FlowSAE: Non-Invasive Sparse Autoencoding with Lifted Invertible Flows

Abstract

Sparse autoencoders (SAEs) are used both to interpret neural network representations and to intervene on them, but a standard SAE's decoder is only an approximate inverse of its encoder, so reconstruction error contaminates any causal effect an intervention is meant to isolate. Prior work compensates for this after the fact, by carrying the error as an external residual channel or by normalising away its damage, rather than removing it. We introduce FlowSAE, to our knowledge the first nonlinear, multi-layer flow-based SAE whose decoder is the exact inverse of its encoder. FlowSAE lifts the activation into a larger space, applies an invertible flow, sparsifies, and inverts. Top- sparsification is the only lossy step. When no feature is removed, the input is recovered exactly for any learned parameters, not only after training. On ImageNet ViT features this holds to numerical precision. It also removes a problem we find in classical SAEs, where uncorrected single-feature ablations can invert the true attribution ranking. At the deployed sparsity, however, exactness does not carry over. The nonlinear flow reconstructs better than a linear exact control, but its feature effects become non-additive, so the standard error-term correction, which is exact for any linear decoder, no longer recovers them. Exact reconstruction adds no asymptotic cost per forward pass, only a constant factor. Making sparse interventions trustworthy is a separate problem.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.