acceptodds
Under review as a conference paper at ICLR 2027

How Does Chain-of-Thought Supervision Teach Models to Reason? A Theoretical Analysis

Abstract

Training with chain-of-thought (CoT) supervision is an effective approach for teaching models multi-step reasoning procedures. At inference time, however, models may need to continue reasoning beyond the finite teacher CoT sequences observed during training, and longer reasoning can instead degrade performance, a phenomenon known as overthinking. How the amount and structure of training-time CoT give rise to such failures remains poorly understood theoretically. To study this problem, we introduce an in-context weight prediction model trained on finite gradient descent trajectories for linear regression. Using high-dimensional asymptotic analysis, we show that the success and failure modes of long-horizon reasoning undergo phase transitions depending on the context length and the total amount of training CoT. We further distinguish two mechanisms of overthinking, error accumulation and error amplification, and show that when the training context is short, a model can learn to continue reasoning before it learns to correct errors. Experiments with fully trained linear and multi-head softmax attention further show that these qualitative behaviors persist. Our results provide a theoretical account linking properties of training-time CoT to inference-time failures and offer guidance for designing teacher CoT that supports stable reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.