acceptodds
Under review as a conference paper at ICLR 2027

Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability

Abstract

Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs through feature sensitivity: the probability that a source-active feature remains active under meaning-preserving paraphrases. Experiments demonstrate that practical scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width×k factorial locates the cause in the active budget k: rare-feature sensitivity declines monotonically as k grows while reconstruction error improves over the same range, so the loss is a property of the selection boundary rather than of width alone. We trace this failure to the geometry of TopK selection: rare active features often lie close to the cutoff between selected and rejected features, so small paraphrase-induced shifts can reorder nearby competitors and remove them from the active set. The distance to that cutoff, the active margin, predicts which features are lost without any threshold, consistently across depths and model families. Guided by the margin diagnosis, we introduce pairwise rank stabilization, which targets the source-paraphrase ordering failure at the cutoff and raises rare-feature sensitivity by 8.83 percentage points on sources that played no role in selecting the objective or its hyperparameters, while reconstruction and alive-feature coverage stay close to the baseline. Overall, our results suggest that wide TopK SAEs should be evaluated not only by reconstruction, sparsity, and feature count, but also by feature reliability under semantic variation and the boundary geometry that determines whether features remain available as stable units of analysis.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.