acceptodds
Under review as a conference paper at ICLR 2027

ExSAL: Explaining Sparse Autoencoder Latents with Domain-Adapted LLMs

Abstract

Sparse Autoencoders (SAEs) decompose language model activations into interpretable features. The standard approach to describing these features prompts a general-purpose LLM with the feature’s top-activating tokens and asks it to produce a natural-language description. However, when the analyzed model has been domain-adapted, its SAE features encode specialized knowledge that a general-purpose explainer may fail to describe accurately. We argue that the domain-adapted model itself is better positioned to explain its own features. To this end, we introduce EXSAL (Explaining Sparse Autoencoder Latents with Domain-Adapted LLMs), a mechanistic interpretability framework for analyzing features from domain-adapted LLMs. Given a domain-adapted model, EXSAL trains an SAE on its activations, and uses that same model to generate natural-language descriptions of the learned features. The framework is applied to two independently adapted models: an enterprise-software model and a finance model. Our experiments show that EXSAL outperforms the general-purpose LLM explainer and narrows the gap to a far larger frontier model, using only the domain-adapted model already at hand and requiring no external explainer. These observations indicate that EXSAL produces more accurate feature descriptions, providing a stronger foundation for analyzing SAE latents in domain-adapted LLMs without requiring a separate explainer model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.