Do We Need Learned Compressors? Adapting LLMs to Parameter-Free KV-Cache Compression for Long-Context Reasoning
Abstract
Large language models (LLMs) face growing memory costs during extended reasoning because the key-value (KV) cache expands with every generated token. Recent methods address this by training auxiliary modules to produce compressed representations of past reasoning states. We ask whether the compression function itself needs to be learned. We propose a parameter-free alternative that periodically replaces fixed-stride windows of reasoning tokens with a single KV pair obtained by simply averaging their pre-rotary key and value representations. Rather than optimizing a learned compressor, we freeze the base model and train lightweight LoRA adapters to reason directly from these fixed summary vectors. Across several long-context reasoning benchmarks and compression ratios, our approach performs competitively with learned KV-cache compression baseline while adding zero auxiliary parameters for compression. Our results suggest that effective reasoning over compressed history does not necessarily require the compression operator itself to be learned.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.