acceptodds
Under review as a conference paper at ICLR 2027

Coverage Is Uniquely Cacheable: Cache-Once Coverage Compression for Multi-Turn Vision-Language Models

Abstract

Vision-language models (VLMs) are increasingly served in multi-turn conversation, where an image is encoded once and its visual key-value (KV) cache is reused across every subsequent turn. Visual token compression, however, is designed for single-pass inference. In multi-turn serving, current methods must either keep the full visual KV resident and re-select tokens for every new question, forfeiting the memory saving, or freeze a first-turn selection that is stale for every later question. We present C³ (Cache-Once Coverage Compression), a training-free method that addresses this challenge. Because the cached subset must serve questions that are unknown at commitment time, C³ abandons query relevance and instead spreads the retained tokens across the image feature space. This enables C³ to preserve a nearby token for whatever visual content a later question may target. The selection depends only on the image, so the cached subset is exactly what any later turn would select, and the whole dialogue runs on one compact visual KV. Across 8 benchmarks, 3 VLMs, and 3 KV budgets, C³ is the only method that avoids severe accuracy collapse across all benchmarks, models, and turns, cutting the persistent visual KV by 6.7x, achieving the lowest cumulative time-to-first-token as dialogues lengthen, and matching or beating the strongest full-KV re-pruning baseline on most benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.