Learning to Carry: KV Cache Eviction for Multi-Turn Vision-Language Models
Abstract
Vision-language models (VLMs) have advanced rapidly in joint visual and textual reasoning, and are now widely applied to multi-turn dialogue. However, their key-value (KV) cache grows with the conversation, increasing memory use and latency. Compressing this cache faces two challenges: eviction must anticipate unknown future queries, and retention must balance modalities despite image entries dominating the cache while receiving little current-turn attention. We propose CarryKV, a learned KV cache eviction method that decides at each turn boundary which entries to carry forward. CarryKV combines Structure-Aware Encoding (SAE), which exposes retention-relevant content and structural cues, with Sketch-Conditioned Selection (SCS), which selects entries complementary to a sketch of randomly retained cache entries. SCS also reserves a minimum image quota before filling the remaining budget by score. We train the selector through future-turn distillation, minimizing the KL divergence between the frozen VLM's predictions with full and retained caches on the following response. Training uses only 100 dialogues, takes under 25 minutes on one A100, and yields a single checkpoint reusable across budgets. Integrated into vLLM, CarryKV compacts the retained cache and physically frees evicted blocks. Across MMDU, MultiVerse, MMRC, and ConvBench on Qwen2.5-VL-7B and GLM-4.6V-Flash-9B, CarryKV maintains strong response quality under aggressive cache compression while reaching - the peak throughput of the full-cache baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.