acceptodds
Under review as a conference paper at ICLR 2027

The Optimizer is an Uncontrolled Variable in Sparse Autoencoder Evaluation

Abstract

Sparse autoencoder (SAE) reconstruction quality is routinely compared across differently trained language models. We show it depends on the optimizer that trained the model, and that the sign of this dependence reverses with the depth at which the SAE is attached. Holding all else fixed, we train 13.9M to 210M-parameter transformers on two corpora under Euclidean (SGD), sign-like (Adam) and spectral (Muon) geometries, with per-geometry learning-rate sweeps and each comparison read separately at matched validation loss and matched training compute (3 to 8 seeds, seed-paired intervals). At the final residual stream and fixed sparsity (TopK, L0 = 32), SAEs fitted to the spectral arm reconstruct worse in every run, by 0.03 to 0.07 explained variance, without attenuation to 210M and surviving four times the training; varying only the update norm reproduces it. Refitting the same SAE at every depth reverses the sign: mid-stack, the spectral arm is reconstructed better, by up to 0.087. The model with the best validation loss is reconstructed worst in every run; identical JumpReLU settings yield a 4.08x spread in achieved sparsity. Cross-model SAE comparisons should report the base optimizer and rate, attachment depth and achieved sparsity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.