acceptodds
Under review as a conference paper at ICLR 2027

Internalizing Chain-of-thought into Looped Transformers

Abstract

The expressive power of transformers concerns both what they can compute and the resources required, including depth, precision, context length, and generated tokens. We organize existing constructions by where they store intermediate states. *In-layer* constructions keep them in hidden activations and produce the answer in one generation step, whereas *in-CoT* constructions record the computation in generated tokens and keep the workspace in the context window. We construct a family of constant-bit-size looped transformers that interpolates between these two classes. For a prescribed loop count and a fixed integer satisfying , our transformer simulates a time-, space- Turing machine using a context window of and generated tokens, reducing both resources by a factor of relative to the corresponding unlooped queue-machine simulation. The construction also holds with the finite-precision softmax kernels specified in this paper.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.