acceptodds
Under review as a conference paper at ICLR 2027

WaveKV: Where Attention Meets Signal Processing for KV-Cache Compression

Abstract

Key-value (KV) caching enables efficient text generation with Large Language Models (LLMs), but its near-linear growth with generated tokens creates increasing memory demands and a memory wall. Existing compression methods rely on query-dependent token eviction or model fine-tuning, which either discard useful KV states or require model adaptation, limiting simultaneous reductions in KV storage and attention computation. Motivated by structured attention patterns and signal-energy redundancy of KV states in the frequency and discrete-wavelet domains, we propose , a training-free KV-cache compression framework that combines discrete cosine transforms (DCT) and discrete wavelet transforms (DWT) to jointly reduce KV-cache storage and attention computation. first performs a one-time, query-independent attention calibration to capture positional attention structure and retrieval-sensitive layers. The resulting positional policy generalizes across datasets and models to guide adaptive DCT-based compression. It then uses adaptive, attention-guided DWT-based hierarchical routing to selectively access the compressed representation during attention computation without directly evicting tokens. Across LLaMA-3.1-8B, Phi3-medium-128k, and Qwen3-8B, achieves \sim\textbf{2\times} KV-cache compression, improving average LongBench and RULER scores over similarly compressed baselines by % and %, respectively, while maintaining competitive language-modeling and long-context capability. At the system level, reduces attention scoring compared to dense full-KV by up to \textbf{11.7\times} and peak memory by % at 64K context, while supporting 128K contexts where full KV runs out of memory. Our code is available at https://anonymous.4open.science/r/WaveKV-95DF

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.