acceptodds
Under review as a conference paper at ICLR 2027

One Full-Attention Layer Is Enough for Hybrid Language Models

Abstract

Hybrid language models replace most softmax attention with linear mixers such as Gated DeltaNet and keep about 25% of the layers as full attention, the recipe of Jamba, Samba, and Qwen3-Next. Those attention layers exist for one capability, long-context retrieval, and they are also the part of the model whose serving cost grows with context: at 32K the six attention layers of a 24-layer 25% hybrid hold 0.38 GB of KV cache per sequence. We ask whether the recipe can be replaced by a hybrid with a *single* full-attention layer. Using 65 pretrained hybrids at 403M and 1.27B (identical data and optimization, 2–6 seeds per cell) and over 80 short continued-pretraining runs, we find that it can, without giving anything up. Against the 25% recipe, one layer matches loss (2.619 vs. 2.616), downstream average, real-data recall (SWDE and FDA), and, after a few hundred needle-augmented continued-pretraining steps, real-text retrieval to 32K, with 1/6 of the attention KV cache (1/24 of full attention). One layer is *necessary*: pure linear models collapse on real-data recall and cannot learn real-text needle retrieval even when trained on it. It is *sufficient* because what bounds real-text retrieval in every architecture we test, from one attention layer to twenty-four, is data and training length: under plain pretraining every model sits on a recency floor beyond 4K, and 500 needle-augmented steps (5% of pretraining tokens) lift a single layer from 0.16 to 0.96 at 8K while 12 or 24 layers learn no better. The remaining choices are secondary knobs with mapped effects: placement (mid-stack for loss and FDA, top for a reliably forming pretrained circuit), looping the layer with shared weights (matches three stacked layers everywhere and beats them when needle data is scarce, at 3× the single-layer KV cache), and inference-time looping (a free unlock for an unformed circuit that becomes harmful once the circuit is trained).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.