FAD-KV: Adaptive Dictionary KV Caching for Efficient Reasoning Decoding
Abstract
Key-value (KV) caching is a core mechanism in large language models, where intermediate key and value representations are stored to accelerate autoregressive decoding. As generation length increases, KV cache memory and computation grow linearly, making compression critical for efficient long-output reasoning. Dictionary-based sparse KV compression is appealing because it represents each token’s KV states as a sparse combination of learned basis elements, or atoms, in a dictionary, preserving token-level structure without enforcing a global low-rank constraint. However, existing dictionary-based methods face two practical limitations during decoding: iterative sparse coding methods such as Orthogonal Matching Pursuit add decoding-time overhead, and static dictionaries may mismatch evolving KV distributions under test-time generation. While adaptive dictionary updates can reduce this mismatch, naive append-only strategies lead to dictionary growth during generation. We propose FAD-KV, a Fast Adaptive Dictionary KV caching framework for efficient reasoning decoding. FAD-KV integrates two components: a screening-based sparse coding scheme, where a one-shot Top-n selection identifies a small candidate pool followed by lightweight coefficient refinement, and a fixed-size dictionary adaptation mechanism that reuses inactive atoms instead of expanding the dictionary. By controlling both sparse coding cost and adaptive dictionary growth, FAD-KV provides a memory-bounded and computation-efficient KV compression framework for long reasoning workloads. Experiments on reasoning and long-context benchmarks show that FAD-KV achieves a strong quality-efficiency trade-off while avoiding unbounded dictionary growth during decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.