Efficient Training of Linear Transformers with Block-diagonal Gates
Abstract
Linear Transformers have established themselves as efficient alternatives to softmax attention for sequence modeling, to the point that modern LLMs increasingly rely on hybrid designs for large-scale deployment. Their primary advantage stems from compressing context into a fixed-size hidden state that is updated recurrently, thus disposing of an ever-growing KV cache. The widespread adoption of Linear Transformers drives the push for their improvement, which hinges on enriching the structure of the state-transition matrix (or *gate*) governing their recurrent update. This structure has progressed from scalar, to diagonal, to diagonal plus a low-rank perturbation: each step broadened the transformations the state could undergo, and was unlocked by dedicated *chunk-wise parallel* algorithms that enabled efficient training at scale. Block-diagonal transitions have recently emerged as the relevant next step, to boost expressivity via richer state mixing. However, the lack of an adequate training algorithm has so far restricted their application to either small scales or to specific block families. In this work, we develop a *generalized chunk-wise parallel* algorithm that lifts these restrictions, together with efficient hardware-aware kernels for its implementation. We use it to introduce and train a Block-diagonal Gated Linear Transformer (BGLT) at 7B scale. To keep parameter count comparable to existing methods, BGLT employs transition blocks, further expressed as a Kronecker product of two factors. Despite this compact representation, BGLT attains superior performance on a range of synthetic tasks, solving in a single layer. On language modeling tasks, it shows results competitive with baselines at 7B. Ultimately, BGLT provides a blueprint for deploying more general recurrences that remain efficiently trainable at scale.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.