Guiding Attention Circuit Discovery with Reasoning Outcomes
Abstract
We adapt attention circuit discovery to reasoning models by learning which attention connections between sentences in a reasoning prefix must be preserved to promote a given reasoning outcome. Because interventions can change the distribution of reasoning continuations, we train masks on continuations sampled from the unmasked model and test them on held-out unmasked continuations and on fresh continuations generated under the learned mask. We consider three objectives—preserving answer distributions, increasing correct-answer probability, and shortening reasoning—and find that our approach is more successful than a local connection baseline that scores connections by their effects on subsequent token predictions. We then investigate the learned circuits and find that retained connections show no consistent preferences for different types of reasoning steps. Moreover, different training seeds also produce distinct circuits with similar performance across sparsity levels. Together, these findings suggest that language model “reasoning” lacks the crisp, systematic structure evoked by the word.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.