acceptodds
Under review as a conference paper at ICLR 2027

Beyond Decoder Proxies: Causal Feature Selection for Sparse Autoencoder Control

Abstract

Scaling sparse autoencoder (SAE) dictionaries creates a feature-selection bottleneck: static decoder projections can screen thousands of candidate directions, but do not reveal which cause a desired probability shift in context. We introduce the target-selection margin (TSM), an intervention-based selection method that ranks features by changes in the log ratio of target to distractor probability mass. Across 18 concepts and five public model/SAE settings, TSM correlates more strongly with held-out margin than decoder proxies (Spearman - versus -). In matched 512-feature pools, paired held-out utility gains have concept-clustered 95% confidence intervals entirely above zero in all five settings. Lexical and grammatical audits broaden coverage beyond the original token sets, while severe target replacement exposes specification sensitivity. Two-dictionary audits identify retrieval losses before causal reranking. Separate behavioral audits show that improved contrast-specific decisions need not overcome full-vocabulary competition and do not establish reliable multi-step generation gains. These results support causal feature selection across prompts under a specified target–distractor contrast, distinct from retrieval coverage and generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.