acceptodds
Under review as a conference paper at ICLR 2027

Local Causal Attribution of Chain-of-Thought Reasoning

Abstract

Understanding the structure of a language model's thought process is an important problem for both transparency and safety. In this work, we take a local and causal approach toward this goal by analyzing the causal relationships among components, termed units, of a given, specific chain-of-thought trace. We construct a structural causal model on these units and relate each unit to the log probability of generating subsequent output units. Our AttriCoT algorithm performs attribution by estimating importance parameters in this structural causal model, using forward passes to compute log probability effects, where is the number of units. Evaluation of perturbation curves across 5 datasets and 4 reasoning models shows that AttriCoT produces attributions that are a more faithful local explanation of the model's behavior than alternative methods. The attribution results also point to differences in thought structure between models and domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.