acceptodds
Under review as a conference paper at ICLR 2027

The Surprising Expressivity of a Single Neuron with Autoregressive Chain-of-Thought

Abstract

The success of autoregressive language models trained with Chain-of-Thought (CoT) supervision raises a fundamental question: How simple can a next-token predictor be while still performing general computation through its autoregressively generated CoT? Surprisingly, we show that a single neuron suffices. For a broad family of activations, including threshold and ReLU, a single autoregressive neuron can simulate any feedforward circuit with the same activation. Specifically, a size- circuit on inputs can be simulated by a neuron of dimension using generation steps, with both bounds independent of circuit depth and tight under natural assumptions. For threshold activations, this improves upon the simulation of Joshi et al. (COLT 2025), whose bounds scale exponentially with circuit depth. Consequently, autoregressive threshold neurons can simulate polynomial-size circuits and hence polynomial-time Turing machines, while remaining efficiently learnable from CoT supervision. Our main expressivity result relies on a novel connection to Sidon sets from additive combinatorics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.