LoGIK: Accelerating SVD approximation with Low-precision Gram Iterative Krylov
Abstract
Low-rank KV cache compression reduces memory by projecting each request's cache onto a compact channel subspace. Adapting this subspace to the realized cache is essential for accurate approximation, but requires online basis extraction, which can become a major bottleneck. Exact SVD provides the optimal rank- subspace but is too costly to run for every request, while faster approximations face their own trade-offs: lightweight methods converge slowly, whereas more accurate Krylov methods repeatedly access the full cache, eroding their speed advantage. We propose Low-precision Gram Iterative Krylov (LoGIK), a Gram-space reformulation of Block Krylov that performs refinement in the much smaller channel space. LoGIK forms the channel Gram matrix once, avoiding subsequent cache traversals and tall token-space orthogonalization. We further accelerate Gram construction with low-precision computation, selectively falling back to FP32 when needed to preserve the retained subspace. This design yields substantial accuracy–latency gains: on a 32k-token Llama-3.1-8B K cache at rank 16, LoGIK matches Block Krylov accuracy while reducing basis-extraction latency by . Integrated into existing KV-compression pipelines, it reduces compression latency by up to while preserving downstream quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.