ExSAL: Explaining Sparse Autoencoder Latents with Domain-Adapted LLMs
Abstract
Sparse Autoencoders (SAEs) decompose language model activations into interpretable features. The standard approach to describing these features prompts a general-purpose LLM with the feature’s top-activating tokens and asks it to produce a natural-language description. However, when the analyzed model has been domain-adapted, its SAE features encode specialized knowledge that a general-purpose explainer may fail to describe accurately. We argue that the domain-adapted model itself is better positioned to explain its own features. To this end, we introduce EXSAL (Explaining Sparse Autoencoder Latents with Domain-Adapted LLMs), a mechanistic interpretability framework for analyzing features from domain-adapted LLMs. Given a domain-adapted model, EXSAL trains an SAE on its activations, and uses that same model to generate natural-language descriptions of the learned features. The framework is applied to two independently adapted models: an enterprise-software model and a finance model. Our experiments show that EXSAL outperforms the general-purpose LLM explainer and narrows the gap to a far larger frontier model, using only the domain-adapted model already at hand and requiring no external explainer. These observations indicate that EXSAL produces more accurate feature descriptions, providing a stronger foundation for analyzing SAE latents in domain-adapted LLMs without requiring a separate explainer model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.