acceptodds
Under review as a conference paper at ICLR 2027

A Comparative Study of In-Context Linear Regression by Transformers with Chain-of-Thought

Abstract

Transformers have demonstrated strong capabilities in multi-step reasoning, with chain-of-thought (CoT) emerging as a powerful prompting and modeling technique. To theoretically understand the capabilities of transformers with CoT, a line of recent work has studied settings in which transformers generate the iterates of an algorithm either through layer-wise propagation, which can be viewed as a form of ”implicit CoT” (Saunshi et al., 2025), or through autoregressive ”explicit CoT” steps, which directly produce the intermediate iterates (Huang et al., 2025a). However, it remains unclear how the transformer and CoT setup affect the model's capabilities. To better understand these mechanisms in a controlled and mathematically concrete setting, we study a single-layer linear self-attention model on two in-context linear regression tasks: weight prediction and query prediction. Using multi-step gradient descent (GD) as the teacher algorithm, we compare when implicit and explicit CoT can or cannot simulate the target iterative computation. Our empirical and theoretical results reveal a sharp contrast in this stylized setting. On weight prediction, implicit CoT fails whereas explicit CoT succeeds; on query prediction, explicit CoT fails, whereas implicit CoT can simulate the GD trajectory up to a fixed output sign transformation and succeeds directly with an added scratchpad. We identify two corresponding bottlenecks: context overwriting during recurrent updates and insufficient writable state for autoregressive token generation. We further show that adding a scratchpad resolves the failure of implicit CoT on weight prediction and enables explicit CoT on query prediction. In addition, we provide a construction showing that the computation realized by explicit CoT can be exactly implemented by implicit CoT with a scratchpad.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.