acceptodds
Under review as a conference paper at ICLR 2027

The Extender: A Log-Structured Transformer

Abstract

We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual , a superposition channel. The Extender adds a concatenation channel : each layer emits both a residual update which is added to , and a much smaller \it extension which is appended to . While both the FFN and see , the attention projections take only as input. As a result, the fully \it extended contains the complete input for the projections of all layers, reducing the persistent attention memory footprint from to . We find that with , the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is smaller than MHA. The memory savings grow with model width.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.