Residual Experts for Sparse Autoencoding
Abstract
Sparse Autoencoders (SAEs) are widely used to extract interpretable features from language models, but scaling full-dictionary SAEs to capture diverse features in large models makes sparse autoencoding computationally challenging. We introduce ResSAE, a two-stage sparse autoencoder that combines conditional computation with residual refinement. A lightweight Level-1 SAE forms an initial reconstruction and uses its active features to select residual experts, small Level-2 SAEs that encode the reconstruction residual. A joint TopK operation allocates the Level-2 sparsity budget across the selected experts. Experiments on Gemma-2-2B residual stream activations show that ResSAE achieves lower reconstruction MSE and higher CE/KL fidelity than representative efficient SAE baselines at matched effective dictionary widths, sparsity levels, and encoder FLOPs. Compared with 16K-feature full-dictionary baselines, ResSAE accesses a 135K-feature space with 56—74% fewer encoder FLOPs while achieving up to 11% lower reconstruction MSE. Controlled ablations show that both residual input and Level-1 routing contribute to the reconstruction gains of residual experts. ResSAE further exhibits an emergent coarse-to-fine feature organization, with reduced feature splitting and absorption while maintaining auto-interpretability comparable to width-matched full-dictionary SAEs. Code and checkpoints will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.