acceptodds
Under review as a conference paper at ICLR 2027

How Large Reasoning Models Solve Problems: A Mechanistic Study on Tower of Hanoi

Abstract

Large reasoning models (LRMs) achieve strong performance that is widely attributed to deliberate thinking in long chains of thought (CoT). Whether the content of these traces is causally relevant to the final answer, however, remains actively debated. In this paper, we present a mechanistic study of CoT traces on a task adapted from the classic Tower of Hanoi puzzle to understand how open-source LRMs solve problems. We selected gpt-oss 20b and gpt-oss 120b, both of which achieve around 95% accuracy on the task under high reasoning effort. We identified 6 dominant strategies the models use to find solutions, including analogy to a known problem, exhaustive search, and means-ends analysis (MEA)–a strategy that has been identified in the study of human planning. Focusing on MEA episodes, we identified 4 dominant reasoning primitives, including subgoal proposal, action proposal, verification, and backtracking. We intervened on these primitives, strengthening or disrupting them by editing the reasoning trace, finding that these interventions had a strong impact on the accuracy and optimality of the resulting solutions. Using causal mediation analysis, we found that verification could be induced by a steering vector built from the outputs of the Mixture-of-Experts (MoE) layers with the strongest causal effects on the model's output. Together, these results provide causal evidence that the content of CoT traces shape the answers generated by LRMs, and show that these traces can be decomposed into meaningful primitives with reliable mechanistic signatures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.