Transformers Provably Learn to Internalize Chain-of-Thought
Abstract
Explicit chain-of-thought (CoT) reasoning improves language models' performance, but generating intermediate tokens increases inference cost. Implicit CoT (ICoT) progressively removes those tokens during training, aiming to internalize reasoning without explicitly generating the intermediate steps at inference. We ask whether a transformer can still learn efficiently as intermediate reasoning steps are progressively removed during training. We introduce Log-ICoT, which uses a given hierarchical decomposition of the task into simpler subproblems. The curriculum removes intermediate steps in geometric chunks aligned with the levels of this hierarchy. We establish sufficient conditions under which a simplified -layer transformer trained with Log-ICoT learns -parity, where , using a number of samples polynomial in the input length and training stages. The resulting model predicts in one forward pass rather than generating a linear number of reasoning tokens sequentially. Experiments illustrate layer-wise internalization on parity and demonstrate the curriculum on balanced ListOps and four-digit multiplication with pretrained GPT-2. Log-ICoT attains high answer accuracy with fewer training updates than vanilla ICoT and lower inference latency than explicit CoT.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.