acceptodds
Under review as a conference paper at ICLR 2027

Concept Attribution Analysis: A Principled Approach to Attributing Changes in LLMs

Abstract

Large language models (LLMs) encode rich semantic information in their internal representations, motivating growing interest in extracting human-interpretable concepts for model understanding and analysis. Existing approaches, including linear probing, sparse autoencoders, and concept component analysis, primarily adopt a single-context perspective, asking what concepts are encoded in the representation associated with a single input context. Such a perspective is inherently underdetermined: multiple concept-level explanations may provide equally plausible interpretations of the same representation, making it difficult to determine which one faithfully reflects the underlying concept structure. In this work, we move from single contexts to pairs of related contexts and exploit the structured variation in their corresponding LLM representations. Building on the recent theoretical characterization of LLM representations through concept log-posteriors, we formulate the representation change between paired contexts as a sum of concept-specific change contributions. We then establish conditions under which these contributions are identifiable up to a global permutation, even when multiple concepts change simultaneously within a given context pair. Guided by this analysis, we develop a practical method, termed Concept Attribution Analysis, that attributes an observed representation change to individual concept changes from paired LLM representations. Experiments on controlled simulations validate the predicted identifiability behavior and demonstrate robustness under departures from the idealized assumptions, while experiments on real LLM representations demonstrate the practical utility of concept-level change attribution.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.