AutoCache: Agentic Search for LLM Prefix Cache Compression
Abstract
Storing and reusing computed states allows LLM serving systems to avoid redundant prefill computation for shared prefixes, reducing time to first token. As context lengths and the number of reusable prefixes grow, however, limited cache capacity forces evictions and expensive recomputation, making storage efficiency critical. This challenge is further complicated by modern hybrid architectures that maintain recurrent and sliding-window states in addition to KV caches, requiring compression strategies that account for model architecture, workload characteristics, and reconstruction cost. We present AutoCache, an agent-driven framework that automatically discovers compression and reconstruction programs for prefix caches. Given a model, dataset, and accuracy target, AutoCache explores a large combinatorial space of compression techniques to compose effective cache-compression recipes, refine existing methods, discover data and model architecture-specific compression opportunities. On the RULER benchmark with 32K context, AutoCache reach 291× KV-cache compression on Qwen3.5-9B, compared with 16.0× for the strongest evaluated baseline. It does so by composing factorization, quantization, entropy encoding, and protection of high-saliency cache entries in ways that would be difficult to design manually. More importantly, AutoCache can seamlessly extend compression beyond KV caches to recurrent states in hybrid models. By discovering and exploiting a previously unexploited low-rank structure in the recurrent state of linear-attention layers, AutoCache achieves 222× compression of the total cache memory footprint. These results demonstrate the potential of automatically discovering compression algorithms that are tailored to the model architecture and workload, therefore enabling substantially more storage-efficient prefix caching for LLM serving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.