acceptodds
Under review as a conference paper at ICLR 2027

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

Abstract

While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, existing techniques necessitate architectural changes to the transformer and incur substantial overhead. In this work, we propose a symmetry-aware value cache merging strategy that reduces cache memory while avoiding both harmful performance degradation and architectural overhead during decoding. Furthermore, we show that this redundancy can be exploited alongside existing cache compression techniques, seamlessly composing with high-ratio quantization and key cache pruning methods. Ultimately, our findings reveal a major source of underutilized capacity in the value cache, opening a simple yet highly effective avenue for scaling context windows under strict memory constraints.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.