acceptodds
Under review as a conference paper at ICLR 2027

SALSA: Training-Consistent Sparse Caching with Sink Preservation for Linearized LLMs

Abstract

The full attention mechanism in standard Transformer architectures incurs quadratic computational complexity with respect to sequence length and linearly growing KV cache overhead, posing severe efficiency bottlenecks in long-sequence scenarios. Although linear attention reduces computational and cache overhead, its compression of history into a fixed-size hidden state prevents precise recall of distant information. Existing conversion methods, such as LoLA, attempt to introduce a sparse cache into hybrid architectures to retain hard-to-compress tokens, but add it only during inference. This causes the attention mechanism learned during conversion to differ from the one actually executed at deployment, leading to a training-inference computation graph mismatch. The core insight of this paper is that the sparse cache should not be merely an inference-time patch; rather, it should be incorporated into the conversion computation graph from the attention distillation stage onward, so that training and inference execute the same attention mechanism. Based on this, we propose SALSA (Sink-Aligned Linearization with Sparse Adaptive Caching), a training-inference-consistent sparse caching linearization framework. SALSA partitions history into three components—a sliding window, a sparse global cache, and a linear attention hidden state—and enables the sparse cache to participate in both attention distillation and fine-tuning, thereby eliminating training-inference distribution shift. On this basis, we further propose two supporting refinements: (1) retaining initial sink tokens in the sliding window to stabilize the attention distribution; and (2) offline layer-wise sparse cache budget allocation based on inter-layer heterogeneity. Experiments on Llama-3.1-8B show that SALSA substantially outperforms prior state-of-the-art methods such as LoLCATs and LoLA on S-NIAH, improving average accuracy by more than 20 absolute points over LoLA while closely matching the full-attention baseline on S-NIAH-1 at 0.5K–2K context lengths. On the efficiency side, SALSA keeps decode cost nearly constant as context grows and, at 32K, achieves higher maximum decode throughput than FA2 under memory constraints, with nearly flat decode memory.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.