DeSKV: Low-Rank Routing and Deferred Selection for KV-Cache Offloading
Abstract
Long-context generation is limited by KV-cache storage and access. Routing-based offloading reduces GPU memory use by keeping exact K/V in host memory and retrieving only a query-dependent working set. This requires a compact router that preserves token rankings as queries change, while minimizing the routing overhead for latency. This paper presents DeSKV, a low-rank routing method with step-deferred selection. The router is fitted from query and key moments, with more weight on keys; Step deferral selection utilizes the preceding step's token selection while preparing the next working set. On Qwen3-8B, DeSKV achieves a 48.3 average LongBench score at a 256-token prompt-KV budget versus 48.5 with full cache, and reduces decode latency by 33.6% relative to full-cache execution at a 97K-token context.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.