Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations
Abstract
Softmax attention is the cornerstone of modern large language models. However, its memory requirements scale linearly with sequence length, and its compute requirements scale quadratically. Linear recurrent models, such as linear attention and state space models, have become widely studied as alternatives to softmax attention due to their linear compute and constant memory requirements. While these sub-quadratic token mixing methods (mixers) achieve promising efficiency gains and competitive results on a wide range of benchmarks, current linear recurrent models still lag behind on tasks that require long-context retrieval or in-context learning. A growing body of work studies hybrid architectures that attempt to mitigate these trade-offs by statically interleaving or merging attention and recurrent blocks. In this work, we explore a new axis of developing hybrid models: across the token sequence. We propose \antelope, a hybrid model that can flexibly switch between different mixers throughout a sequence, e.g., quadratic attention for rich context utilization and linear recurrences to reduce quadratic computation. Oryx ties at least 90% of its parameters across mixers, enabling attention and recurrent modes to operate over shared internal representations. We validate our design with Mamba-2 and Gated DeltaNet variants, up to 1.4B models. Under fixed token budgets and a mixed-training strategy, Oryx achieves comparable or better performance than its single-mixer baselines. At the 1.4B scale, all instances of Oryx outperform their respective baselines by at least 0.7 percentage points on averaged language modeling tasks. On retrieval tasks, Oryx achieves performance comparable to the Transformer baseline even when processing only a tiny fraction (<10%) of the tokens in the attention mode. These results suggest that attention and linear recurrent models can share internal representations, and motivate sequence-axis hybridization as a promising direction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.