Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability
Abstract
Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs through feature sensitivity: the probability that a source-active feature remains active under meaning-preserving paraphrases. Experiments demonstrate that practical scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width×k factorial locates the cause in the active budget k: rare-feature sensitivity declines monotonically as k grows while reconstruction error improves over the same range, so the loss is a property of the selection boundary rather than of width alone. We trace this failure to the geometry of TopK selection: rare active features often lie close to the cutoff between selected and rejected features, so small paraphrase-induced shifts can reorder nearby competitors and remove them from the active set. The distance to that cutoff, the active margin, predicts which features are lost without any threshold, consistently across depths and model families. Guided by the margin diagnosis, we introduce pairwise rank stabilization, which targets the source-paraphrase ordering failure at the cutoff and raises rare-feature sensitivity by 8.83 percentage points on sources that played no role in selecting the objective or its hyperparameters, while reconstruction and alive-feature coverage stay close to the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and the boundary geometry that determines whether features remain available as stable units of analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.