acceptodds
Under review as a conference paper at ICLR 2027

InvExplain: An Agentic Interpretation Framework for SAE Features via Training-Free Diffusion Inversion

Abstract

Large language models (LLMs) have achieved remarkable capabilities, yet their internal representations remain difficult to understand. Sparse autoencoders (SAEs) offer a promising approach by decomposing these representations into more interpretable features. However, automatically interpreting individual SAE features remains challenging. Existing automated pipelines primarily infer explanations from corpus-retrieved activating samples, limiting their evidence to activation patterns represented in a fixed corpus. Feature inversion can expand evidence beyond the corpus, but learned inverse models require additional training and remain tied to their training distributions. To overcome these limitations, we introduce InvExplain, an agentic interpretation framework driven by training-free diffusion posterior sampling (DPS). By leveraging a frozen diffusion language model prior, InvExplain actively acquires new, activation-verified evidence beyond the retrieval corpus without additional training. An LLM agent combines generated and corpus-retrieved evidence to formulate and test competing hypotheses, uses DPS to repair failed positive tests, and refines the explanation. Theoretically, we characterize the prior-dependent distortion of training-based inversion and show that diffusion posterior sampling can be viewed as approximating an activation-constrained distribution under a pretrained language prior. Across three models, InvExplain outperforms corpus-reliant baselines on a shared test of whether a feature activates on a text and, if so, where its activation peaks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.