PiCoKV: Dequantization-Free KV Cache Quantization via Binary Principal-Component Coding
Abstract
The growing key–value (KV) cache limits long-context large language model (LLM) inference through its memory footprint and bandwidth demand. Quantization reduces these costs, but attention typically still requires dequantizing or unpacking the stored KV states into wider numerical operands for computation, introducing conversion overhead and limiting the computational benefits of low-bit storage. To bridge this gap between compact storage and computation, we present PiCoKV, a training-free, data-aware KV-cache quantization method for dequantization-free attention on packed binary codes. PiCoKV factorizes keys and values into binary codes and compact decoding matrices. By folding these matrices into attention's query and output paths, PiCoKV performs both query–key scoring and value aggregation directly on packed codes through word-level bitwise operations, processing multiple binary entries in parallel without unpacking them into wider numerical operands. Evaluations on long-context benchmarks for Llama-3.1-8B-Instruct and Qwen3-8B show that PiCoKV achieves accuracy comparable to full attention and outperforms low-bit quantization baselines. The combination of compact storage and direct computation yields practical efficiency gains: PiCoKV achieves higher end-to-end decoding throughput than the baselines, reaching that of FlashAttention-2 at a 256K context while compressing the KV cache by up to . Our code is available at https://anonymous.4open.science/r/PiCoKV-BFC1/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.