A Comparative Study of In-Context Linear Regression by Transformers with Chain-of-Thought
Abstract
Transformers have demonstrated strong capabilities in multi-step reasoning, with chain-of-thought (CoT) emerging as a powerful prompting and modeling technique. To theoretically understand the capabilities of transformers with CoT, a line of recent work has studied settings in which transformers generate the iterates of an algorithm either through layer-wise propagation, which can be viewed as a form of ”implicit CoT” (Saunshi et al., 2025), or through autoregressive ”explicit CoT” steps, which directly produce the intermediate iterates (Huang et al., 2025a). However, it remains unclear how the transformer and CoT setup affect the model's capabilities. To better understand these mechanisms in a controlled and mathematically concrete setting, we study a single-layer linear self-attention model on two in-context linear regression tasks: weight prediction and query prediction. Using multi-step gradient descent (GD) as the teacher algorithm, we compare when implicit and explicit CoT can or cannot simulate the target iterative computation. Our empirical and theoretical results reveal a sharp contrast in this stylized setting. On weight prediction, implicit CoT fails whereas explicit CoT succeeds; on query prediction, explicit CoT fails, whereas implicit CoT can simulate the GD trajectory up to a fixed output sign transformation and succeeds directly with an added scratchpad. We identify two corresponding bottlenecks: context overwriting during recurrent updates and insufficient writable state for autoregressive token generation. We further show that adding a scratchpad resolves the failure of implicit CoT on weight prediction and enables explicit CoT on query prediction. In addition, we provide a construction showing that the computation realized by explicit CoT can be exactly implemented by implicit CoT with a scratchpad.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.