acceptodds
Under review as a conference paper at ICLR 2027

Sparse Autoencoders vs. Transcoders: Encoder Direction Use and Single-Feature Discrimination

Abstract

Sparse autoencoders (SAEs) reconstruct MLP outputs, whereas transcoders (TCs) predict them from the corresponding MLP inputs. We compare their single-feature discrimination and examine how encoder direction use and JumpReLU contribute to the difference. Across all 26 layers of Gemma 3 1B, we construct shared MLP-output clusters, select each coder's best feature by postactivation AUROC, and evaluate on held-out documents. TC achieves higher postactivation AUROC in 24 of 26 layers under both sparsity settings, despite SAE having higher mean preactivation AUROC. JumpReLU maps distinct preactivation values to zero, producing a larger discrimination loss for SAE. In an exploratory intervention, reducing TC's reliance on input principal components using an SAE-derived reference narrows its postactivation advantage more than matched random rotations, identifying encoder weighting of high-variance input directions as a partial contributor. Separately, TC retrieves approximately three and four additional target sentences per 200 reviewed under the lower- and higher-activity settings, respectively. Together, these results link single-feature discrimination to both encoder direction use and learned activation thresholds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.