Non-Gaussianity Reveals Interpretable Directions in Language Models
Abstract
Finding human-interpretable directions in language-model activations is important for understanding and controlling model behavior. These directions are usually discovered through indirect learning objectives such as sparse reconstruction, but it remains unclear what existing statistical structure distinguishes interpretable directions and whether that structure can be used directly for feature discovery. We find that non-Gaussianity provides such a signature: projection distributions of natural-text residual-stream activations along sparse autoencoder (SAE) decoder directions are substantially more non-Gaussian than those along random directions, with most contexts forming a concentrated background and recurring patterns often occupying one tail. This raises a natural question: if interpretable directions tend to have non-Gaussian projection distributions, can we discover them by searching for non-Gaussian directions directly in activation space, without training another sparse dictionary? To test this hypothesis, we introduce ICA Lens, a practical workflow for efficient discovery, reliable annotation, and faithful evaluation of non-Gaussian directions in LLM residual-stream activations. Across GPT-2 Small, Gemma 2 2B, and Qwen 3.5 9B Base, we find that: (1) the recovered non-Gaussian directions capture coherent lexical, syntactic, and semantic patterns under both human and LLM-based evaluation; (2) non-Gaussianity is positively associated with both human-label confidence and automated interpretability; and (3) the resulting coordinates are competitive with SAEs across interpretability, sparse reconstruction, concept prediction, and targeted probe perturbation, while fitting in minutes per layer and using 7–43× less memory. These results show that non-Gaussianity is not merely a statistical property of learned features, but a direct unsupervised signal for discovering interpretable and useful structure in activations. We release our full implementation, including code to reproduce all experiments and the complete ICA Lens workflow, in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.