The Context-Ready Transformer
Abstract
We introduce the *context-ready transformer*, a decoder architecture that pre-contextualizes each token before it enters the transformer block. During left-to-right generation, a correction network combines the previous position’s final block representation with the current token embedding. The token enters the block as its embedding plus this correction, while ordinary causal attention and the standard key-value (KV) cache remain unchanged. The central computation separates training depth from online inference depth. A context-ready model has a \(D\)-layer block. Parallel training performs \(K\) shared block evaluations over the full sequence, giving effective training depth \(DK\). Streaming inference processes tokens left to right with one block evaluation per new token and online per-token critical path \(D\). A pretrained decoder-only transformer can also be converted by adding a zero-initialized correction feed-forward network (FFN) and fine-tuning. We evaluate validation perplexity (PPL) across widths, depths, block sizes, and two datasets, with comparisons against standard transformers and ablations. In single-run comparisons, a \(D=5\) model obtains lower PPL than a 12-layer transformer while generating \(1.7\times\) faster on an A100. Using \(K=10\) during training, a single-layer model (\(D=1\)) obtains lower PPL than a 6-layer transformer with a \(2.6\times\) inference speedup. Parallel training therefore requires \(K\) block evaluations per iteration, whereas streaming inference requires one per generated token. For the checkpoints in Table 4, sequential inference and parallel \(K=10\) evaluation differ by at most 0.01 PPL. Separately, in a 10-hop pointer-chasing task that requires backpropagation through time (BPTT) across the full correction chain, a \(D=1\) model solves all 11 evaluated query groups, while standard transformers exhibit staircase-like depth dependence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.