Beyond Decoder Proxies: Causal Feature Selection for Sparse Autoencoder Control
Abstract
Scaling sparse autoencoder (SAE) dictionaries creates a feature-selection bottleneck: static decoder projections can screen thousands of candidate directions, but do not reveal which cause a desired probability shift in context. We introduce the target-selection margin (TSM), an intervention-based selection method that ranks features by changes in the log ratio of target to distractor probability mass. Across 18 concepts and five public model/SAE settings, TSM correlates more strongly with held-out margin than decoder proxies (Spearman - versus -). In matched 512-feature pools, paired held-out utility gains have concept-clustered 95% confidence intervals entirely above zero in all five settings. Lexical and grammatical audits broaden coverage beyond the original token sets, while severe target replacement exposes specification sensitivity. Two-dictionary audits identify retrieval losses before causal reranking. Separate behavioral audits show that improved contrast-specific decisions need not overcome full-vocabulary competition and do not establish reliable multi-step generation gains. These results support causal feature selection across prompts under a specified target–distractor contrast, distinct from retrieval coverage and generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.