Identifying Changes in Causal Attribution under Reasoning Step Replacement
Abstract
Language models can replace written reasoning steps with continuous hidden states. How does this change the contribution of the reasoning that remains visible? We measure how a change in the input affects the answer by intervening on the states stored for the prompt, written steps, and hidden steps. These parts interact, so assigning credit to one depends on how their joint effects are divided. We compare two reasoning modes using the same rule for assigning credit. For rules based on different orderings of the parts, we derive the exact range of possible changes in attribution. Unchanged interactions cancel; only changes in interactions make this comparison ambiguous. We apply the method to GPT-2 models trained on a synthetic policy task, holding each model's weights fixed across modes. Across three adaptation runs, replacing more steps increases the contribution of states at text positions, while accuracy and correct policy use decline. None of the runs supports an increase in the contribution of hidden-step states. In matched comparisons within two models, interventions on text inputs produce smaller attribution changes than interventions on the stored states. The results separate the influence carried by text-associated states from the model's ability to apply the policy correctly.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.