Concept Component Analysis: A Principled Approach for Concept Extraction in LLMs
Abstract
Extracting human-interpretable concepts from LLMs is essential to understand the mechanisms underlying their capabilities. Sparse autoencoders (SAEs) approach this task by decomposing LLM internal representations into a sparse dictionary, with the hope that the extracted features align with human-interpretable concepts. Despite their empirical progress, SAEs suffer from a fundamental theoretical ambiguity: the relationship between LLM representations and human-interpretable concepts remains unclear. This lack of theoretical grounding gives rise to several methodological challenges, especially how sparsity should be imposed to facilitate concept extraction. In this work, we extend recent theoretical analyses of LLM representations and show that LLM representations can be approximated as a linear mixture of the marginal log-posteriors of individual latent concepts conditioned on the input context. Importantly, this characterization extends beyond the final layer to intermediate layers of LLMs. These findings motivate CONcept Component Analysis (ConCa), a principled formulation of concept extraction as a linear unmixing problem for recovering the log-posterior coordinates of individual concepts. We develop a sparse instantiation of ConCa, termed Sparse ConCa, and demonstrate its ability to extract meaningful concepts across multiple LLMs, offering theory-guided advantages over existing SAEs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.