acceptodds
Under review as a conference paper at ICLR 2027

Multi-Mixer Models: Flexible Sequence Modeling with Shared Representations

Abstract

Softmax attention is the cornerstone of modern large language models. However, its memory requirements scale linearly with sequence length, and its compute requirements scale quadratically. Linear recurrent models, such as linear attention and state space models, have become widely studied as alternatives to softmax attention due to their linear compute and constant memory requirements. While these sub-quadratic token mixing methods (mixers) achieve promising efficiency gains and competitive results on a wide range of benchmarks, current linear recurrent models still lag behind on tasks that require long-context retrieval or in-context learning. A growing body of work studies hybrid architectures that attempt to mitigate these trade-offs by statically interleaving or merging attention and recurrent blocks. In this work, we explore a new axis of developing hybrid models: across the token sequence. We propose \antelope, a hybrid model that can flexibly switch between different mixers throughout a sequence, e.g., quadratic attention for rich context utilization and linear recurrences to reduce quadratic computation. Oryx ties at least 90% of its parameters across mixers, enabling attention and recurrent modes to operate over shared internal representations. We validate our design with Mamba-2 and Gated DeltaNet variants, up to 1.4B models. Under fixed token budgets and a mixed-training strategy, Oryx achieves comparable or better performance than its single-mixer baselines. At the 1.4B scale, all instances of Oryx outperform their respective baselines by at least 0.7 percentage points on averaged language modeling tasks. On retrieval tasks, Oryx achieves performance comparable to the Transformer baseline even when processing only a tiny fraction (<10%) of the tokens in the attention mode. These results suggest that attention and linear recurrent models can share internal representations, and motivate sequence-axis hybridization as a promising direction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.