acceptodds
Under review as a conference paper at ICLR 2027

Pipelined Transformer: Autoregressive Generation with Increased Parallelism

Abstract

Transformers are slow at generating text. Each token must be *fully* generated *before* work starts on the next token. Diffusion models break this serial decoding bottleneck, but they are non-autoregressive latent-variable models. We propose **PipeT**, a *pipelined* architecture that remains autoregressive and simple to train. Tokens are decoded in parallel while attending to one another bidirectionally, as in diffusion, but token is slightly delayed with respect to token , so it can depend autoregressively on token and all previous tokens. This architecture decodes faster than a standard Transformer, reduces KV cache size, and is formally more expressive. Although pretraining and prompt prefilling are slower than in a standard Transformer, this tradeoff may be worth it when generation is a bottleneck, as when generating code or long reasoning traces during RL training or after deployment. We evaluate models with 4M–200M parameters trained on the SlimPajama-6B subset using Chinchilla-style token budgets. At 100M parameters, a 3-layer **PipeT** matches the test loss of a 4-layer standard Transformer of the same size while decoding faster with a smaller KV cache. It nearly matches a 16-layer standard Transformer (the optimal number of layers that we evaluated at this parameter budget), with only nats/token higher test loss while decoding faster with a smaller KV cache.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.