DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
Abstract
The increasing memory footprint of the KV cache and the high computational cost of query-key similarity computation constitute two fundamental bottlenecks for large language models in long-context inference. While existing KV cache compression methods primarily alleviate the former, they often sacrifice generation quality and leave the latter largely unaddressed. This paper introduces DASH-KV, an innovative acceleration framework that reformulates attention as approximate nearest neighbor search via asymmetric deep hashing. Under this paradigm, we design an asymmetric encoding architecture that differentially maps queries and keys to account for their distinctions in precision and reuse characteristics. To balance efficiency and accuracy, we further introduce a dynamic mixed-precision mechanism that adaptively retains full-precision computation for critical tokens. Experiments on six LongBench tasks across three backbones show that DASH-KV achieves performance comparable to full attention. At a 128K context length, DASH-KV achieves a 2.732x speedup over pre-expanded SDPA for the complete single-layer decoding attention path on Qwen2-7B-Instruct.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.