RouteCache: Efficient Long-Context Attention via Learnable Retrieval and Temporal Caching
Abstract
Sparse attention can reduce long-context inference costs, but learning discrete token selection from the language modeling objective remains challenging. We introduce RouteCache, which combines end-to-end trained hash-based retrieval with an LRU working cache for efficient decoding. Our key insight is that attention-score gradients provide task-driven supervision for retrieval scores, enabling training beyond layer-wise imitation of dense attention. Starting from oracle initialization, we introduce closed-loop gradient feedback: gated hash scores influence attention during training and receive feedback from the resulting language modeling loss. This jointly optimizes hash networks across layers while keeping the pretrained backbone frozen. At inference, the gates are removed and the hash networks select tokens for sparse attention across all layers. To translate sparse retrieval into practical acceleration, the working cache reuses recently retrieved KV pairs, while fused kernels and asynchronous execution reduce data movement and scheduling overhead. Across Qwen3 models from 4B to 30B parameters, including MoE and reasoning variants, RouteCache achieves quality generally close to dense attention on perplexity, LongBench V2, and RULER evaluations. At 256K context, latency benchmarks show up to 4.9–7x speedups over StaticCache, depending on the retrieval and cache budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.