acceptodds
Under review as a conference paper at ICLR 2027

On-Policy Attention Linearization

Abstract

Hybrid transformer architectures that replace most softmax attention layers with linear attention offer transformer-level quality at a fraction of the memory cost. Rather than pretraining such models, a growing body of work distills them from already trained full-attention transformers by aligning the attention layers' hidden states and matching the models' next-token distributions. We find that these distilled models often collapse on long-context retrieval and reasoning tasks, particularly when operating in thinking mode, where the efficiency gains of hybrid architectures matter most. Since linear attention layers must compress context into a fixed-size state, their errors compound over long sequences. As off-policy distillation never teaches the student model to recover from this drift, tasks that necessitate longer sequence lengths become especially challenging. We introduce On-Policy Attention Linearization (OPAL) in which the hybrid attention student samples its own long-context trajectories and receives dense supervision from the frozen full-attention teacher. Applying OPAL to Qwen3-4B and MiMo-7B-RL-0530, we recover of full attention performance on commonsense reasoning, on needle-in-a-haystack (NIAH) retrieval, and on mathematical reasoning with only 3B training tokens. We achieve these results without Supervised Fine-Tuning (SFT) or Reinforcement Learning with Verifiable Rewards (RLVR) training stages and outperform competitive methods that utilize such techniques by on long-context retrieval in terms of recovery to teacher.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.