acceptodds
Under review as a conference paper at ICLR 2027

HayateQuant: More Geometry per Bit with Folded Rotations for KV Cache Quantization

Abstract

Compressing the key–value (KV) cache can substantially reduce memory traffic during language-model decoding. However, improved quantization geometry is useful only if it can be realized with low encoding and serving overhead. We investigate whether short-block quantization can improve a data-free cache codec while preserving a compact, fixed-rate representation. We introduce HayateQuant (HQ), which replaces independent scalar assignments with compact joint codebooks. A folded rotation promotes uniform within-block directions, while precomputed candidate lists keep online assignment efficient. To account for the distinct effects of key and value quantization errors, HQ further employs role-specific scaling and key and value bit allocation. Across diverse model families, HayateQuant achieves a favorable quality-compression tradeoff, particularly at moderate rates. Its optimized packed serving implementation also improves decoding throughput over the existing baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.