acceptodds
Under review as a conference paper at ICLR 2027

DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing

Abstract

The increasing memory footprint of the KV cache and the high computational cost of query-key similarity computation constitute two fundamental bottlenecks for large language models in long-context inference. While existing KV cache compression methods primarily alleviate the former, they often sacrifice generation quality and leave the latter largely unaddressed. This paper introduces DASH-KV, an innovative acceleration framework that reformulates attention as approximate nearest neighbor search via asymmetric deep hashing. Under this paradigm, we design an asymmetric encoding architecture that differentially maps queries and keys to account for their distinctions in precision and reuse characteristics. To balance efficiency and accuracy, we further introduce a dynamic mixed-precision mechanism that adaptively retains full-precision computation for critical tokens. Experiments on six LongBench tasks across three backbones show that DASH-KV achieves performance comparable to full attention. At a 128K context length, DASH-KV achieves a 2.732x speedup over pre-expanded SDPA for the complete single-layer decoding attention path on Qwen2-7B-Instruct.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.