acceptodds
Under review as a conference paper at ICLR 2027

Readout Scaling Shapes Sparse Feature Recovery

Abstract

Sparse autoencoders (SAEs) can learn different features from representations that contain the same information. We study this effect at a neural network’s linear readout, whose singular-value gains change the reconstruction weights seen by an SAE. Readout-canonicalized decision-space (RC-DS) coordinates remove these gains while preserving retained decision information. Controlled synthetic studies show that the generating distribution determines which coordinates recover known factors best. When factors are generated before readout scaling, removing gains at condition number 100 raises mean recovery from 0.5063 to 0.9134. Changing the generating coordinates reverses the preferred method. An exact wrapper reproduces RC-DS training to numerical precision, isolating training geometry rather than additional information. In bargaining and Goofspiel policies, RC-DS improves mean logit reconstruction under fixed budgets. On three frozen language-model checkpoints, it improves gain-free reconstruction in all 18 tested task–sparsity setting means, while logit training better reconstructs output logits in 17. Longer training and separate tuning reduce gaps on small label heads. Label-detection and steering controls show that reconstruction gains do not establish more useful features. Together, these studies identify how readout scaling, data geometry, and training determine the effect of an information-preserving coordinate change.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.