acceptodds
Under review as a conference paper at ICLR 2027

RouteCache: Efficient Long-Context Attention via Learnable Retrieval and Temporal Caching

Abstract

Sparse attention can reduce long-context inference costs, but learning discrete token selection from the language modeling objective remains challenging. We introduce RouteCache, which combines end-to-end trained hash-based retrieval with an LRU working cache for efficient decoding. Our key insight is that attention-score gradients provide task-driven supervision for retrieval scores, enabling training beyond layer-wise imitation of dense attention. Starting from oracle initialization, we introduce closed-loop gradient feedback: gated hash scores influence attention during training and receive feedback from the resulting language modeling loss. This jointly optimizes hash networks across layers while keeping the pretrained backbone frozen. At inference, the gates are removed and the hash networks select tokens for sparse attention across all layers. To translate sparse retrieval into practical acceleration, the working cache reuses recently retrieved KV pairs, while fused kernels and asynchronous execution reduce data movement and scheduling overhead. Across Qwen3 models from 4B to 30B parameters, including MoE and reasoning variants, RouteCache achieves quality generally close to dense attention on perplexity, LongBench V2, and RULER evaluations. At 256K context, latency benchmarks show up to 4.9–7x speedups over StaticCache, depending on the retrieval and cache budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.