acceptodds
Under review as a conference paper at ICLR 2027

Keep Features Dense and Explanations Sparse: Adaptive Interpretability at Test Time

Abstract

Simple data-derived linear directions can expose highly interpretable structure in neural representations, yet their activations are typically signed and dense. Sparse autoencoders instead learn non-negative sparse features, coupling feature learning with a sparsity constraint. We ask whether sparsity is needed in the representation at all. More broadly, test-time computation remains comparatively underexplored in interpretability. We introduce **TTS-AE**, which learns a dense, non-negative reconstructive representation from data-grounded directions and introduces sparsity only when an explanation is requested. At test time, *Test-Time Sparsification* (TTS) selects a compact, input-specific subset of features and infers the remaining activations from the covariance structure of the dense representation until a desired reconstruction fidelity is reached. Across three model–dataset pairs, the resulting dense features are more interpretable than nonlinear sparse autoencoder features, while TTS reaches matched fidelity with substantially fewer positively activated concepts than sparse autoencoders and sparse approximation baselines. The same dense concept space also supports covariance-aware interventions and sparse target-specific global explanations. Our results suggest that interpretability and explanation sparsity need not be learned jointly: dense representations can support sparse explanations constructed adaptively through additional computation at test time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.