How Does Chain-of-Thought Supervision Teach Models to Reason? A Theoretical Analysis
Abstract
Training with chain-of-thought (CoT) supervision is an effective approach for teaching models multi-step reasoning procedures. At inference time, however, models may need to continue reasoning beyond the finite teacher CoT sequences observed during training, and longer reasoning can instead degrade performance, a phenomenon known as overthinking. How the amount and structure of training-time CoT give rise to such failures remains poorly understood theoretically. To study this problem, we introduce an in-context weight prediction model trained on finite gradient descent trajectories for linear regression. Using high-dimensional asymptotic analysis, we show that the success and failure modes of long-horizon reasoning undergo phase transitions depending on the context length and the total amount of training CoT. We further distinguish two mechanisms of overthinking, error accumulation and error amplification, and show that when the training context is short, a model can learn to continue reasoning before it learns to correct errors. Experiments with fully trained linear and multi-head softmax attention further show that these qualitative behaviors persist. Our results provide a theoretical account linking properties of training-time CoT to inference-time failures and offer guidance for designing teacher CoT that supports stable reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.