Local Causal Attribution of Chain-of-Thought Reasoning
Abstract
Understanding the structure of a language model's thought process is an important problem for both transparency and safety. In this work, we take a local and causal approach toward this goal by analyzing the causal relationships among components, termed units, of a given, specific chain-of-thought trace. We construct a structural causal model on these units and relate each unit to the log probability of generating subsequent output units. Our AttriCoT algorithm performs attribution by estimating importance parameters in this structural causal model, using forward passes to compute log probability effects, where is the number of units. Evaluation of perturbation curves across 5 datasets and 4 reasoning models shows that AttriCoT produces attributions that are a more faithful local explanation of the model's behavior than alternative methods. The attribution results also point to differences in thought structure between models and domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.