Gated Recurrent Transformers: Expressive Depth Through Recurrent Modulation
Abstract
Scaling transformer language models creates an inherent tension between expressivity and memory efficiency. While unique weights across layers preserve functional specialization—from input-grounding to abstract refinement—they incur a substantial memory footprint. Conversely, standard depth-sharing enforces uniform transformations that collapse representational diversity and degrade modeling quality. We introduce the GATED RECURRENT TRANSFORMER (GRT), a recurrent depth transformer inspired by gated recurrent neural networks, where a lightweight projection and an elementwise update gate—conditioned on the hidden state, a fixed prelude output, and noise resampled at every step—modulate the recurrent update. This allows the model to specialize the input to the same few layers across recurrences, rather than requiring many unique layers to achieve functional diversity. Under an isoFLOPS constraint, a 3-layer GRT matches the accuracy of a 12-layer GPT-2 Small baseline with similar training and inference FLOPs, and leads a wide range of existing recurrent depth approaches in all nine scale-by-budget cells; at medium and large scale it approaches dense quality at the standard token budget and overtakes it at medium scale once that budget is doubled. Under an isoPARAMS constraint, deeper recurrence achieves a 2.76 validation loss versus 2.84 for a non-recurrent counterpart at matched parameter and data budget. Our results demonstrate that adaptive depth reuse is a principled strategy for trading parameters for quality. Intermediate recurrences additionally support lossless greedy self-speculative decoding without a separate draft model. Moreover, we propose an optional lightweight post-training optimization, Channel-wise KV Codec, to compress recurrent KV caches to the storage budget of a single K/V pair per token.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.