FastRebase: Hybrid Low-Rank States for Efficient Linear-Attention Decoding
Abstract
Linear attention replaces the growing key–value cache of Transformers with a fixed-size recurrent state, but repeatedly accessing this state can still incur substantial memory traffic during batched decoding. Low-rank approximation offers a natural way to reduce this cost. However, maintaining such representations online is challenging due to the cost of repeated singular value decompositions (SVDs). To address this, we introduce FastRebase, which maintains the recurrent states of linear attention models using a hybrid low-rank representation throughout inference without online SVDs. Our key insight is that rank allocation and subspace tracking can be decoupled: FastRebase determines a per-head rank profile offline and tracks only a sequence-dependent subspace during inference using GPU-efficient projections and small-matrix solves. Its hybrid state combines a low-rank history with a residual buffer of recent key–value associations, enabling periodic rebasing that amortizes the cost of subspace updates over multiple decoding steps. We implement FastRebase with custom Triton kernels and evaluate it on models with scalar or diagonal decay and delta-rule state transitions. FastRebase reduces recurrent-state traffic by 61.7–87.3%, while closely matching the dense baselines in zero-shot accuracy and long-context perplexity. At batch size 128 on an H100 GPU, the traffic reduction translates into 1.36–4.73× speedups for linear-attention recurrence and 1.18–2.31× for end-to-end decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.