acceptodds
Under review as a conference paper at ICLR 2027

Rational Sparse Autoencoder

Abstract

Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder nonlinearities such as ReLU, JumpReLU, and TopK. This hard-codes a particular sparsity mechanism into the model and can distort the reconstruction-versus-sparsity trade-off. We introduce the *Rational Sparse Autoencoder* (RSAE), which replaces the fixed encoder activation with a trainable rational function. Our idea is developed based on a theoretical reconstruction advantage for RSAE at matched expected sparsity, with guaranteed non-inferiority to ReLU and, under stated assumptions, provable improvements over JumpReLU and TopK. The model is then implemented through a two-stage pipeline: an initialisation procedure that copies the pre-trained baseline SAE weights, and calibrates the scale parameters along with the rational coefficients; followed by a fine-tuning step under a sparsity-regularised reconstruction objective that holds at the teacher's value. Consistent with this theory, on residual-stream activations of three open-weight language models and across all three baseline activation families, the RSAE *improves* reconstruction fidelity after the fine-tuning step at the same level of sparsity. These gains are consistent across host language models, across baseline activation families, and across the full range of baseline sparsity we tested, while the upgrade itself adds only 20 scalar parameters per autoencoder and one training run takes between about 4 and 25 minutes on a single consumer GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.