Guided Sparse Autoencoders for Interpreting Reasoning Features
Abstract
Large Language Models (LLMs) increasingly solve complex tasks through multi-step Chain-of-Thought (CoT) reasoning, yet the internal representations supporting these trajectories remain difficult to interpret. Sparse Autoencoders (SAEs) decompose dense activations into sparse latent features, but standard reconstruction-sparsity only training often misses subtle reasoning-related features because they contribute negligibly to overall activation reconstruction variance. Consequently, important reasoning primitives do not naturally occupy identifiable latent directions. We propose the Guided Sparse Autoencoder (G-SAE), a partitioned latent-space architecture that injects supervision directly into sparse representation learning. A subset of latent dimensions is jointly supervised for step reasoning correctness, while the remaining dimensions retain an unsupervised role in reconstructing broader activation information. By combining labeled ProcessBench activations with unlabeled activations, G-SAE merges targeted correctness pressure with broad reconstruction coverage. Evaluations across different model scales demonstrate that supervised latent partitioning improves correctness prediction while maintaining reconstruction fidelity.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.