ROKI: RoPE-Aware KV Cache Quantization via Isofrequency Grouping
Abstract
Low-bit KV-cache quantization reduces the memory and bandwidth costs of long-context LLM inference, but post-RoPE per-token quantization is sensitive to large activations that dominate the shared quantization range. Cross-channel rotations can reduce these outliers, yet general cross-pair rotations do not commute with Rotary Position Embedding (RoPE) and therefore cannot be folded into Q/K projection weights. We propose ROKI, a RoPE-aware KV-cache quantization framework based on Isofrequency Commutant Expansion (ICE). Rather than restricting transformations to the fixed RoPE operator, ICE selectively ties nearby low-frequency RoPE pairs, creating isofrequency groups that enlarge the exact RoPE-commuting transformation space from pair-wise to cross-pair rotations. ROKI uses this expanded space to redistribute Q/K activations and applies transformation-aware mixed-precision key quantization based on post-transformation QK contributions. All Q/K transformations are optimized offline and folded into the projection weights, requiring no additional online transformation. At an average 3-bit KV-cache budget, ROKI achieves the highest average LongBench-E score among the evaluated quantized baselines on three Llama/Qwen models while remaining close to FP16. Our ROKI implementation also achieves up to 4.22 decode speedup over the evaluated low-bit baselines at long contexts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.