acceptodds
Under review as a conference paper at ICLR 2027

Query-Guided Coarse-to-Fine KV Cache Compression for Unified Multimodal Generation

Abstract

Unified multimodal models have demonstrated strong capabilities in visual generation and editing, but long multimodal KV caches incur substantial memory and attention costs during generation. We propose a coarse-to-fine KV cache compression framework that progressively prunes redundant KV tokens while controlling the cost of importance estimation. The coarse stage exploits the local similarity of visual queries and their similar preferences over cached tokens. We sample one representative query from each local group to reduce the QK computation required for importance estimation. For stronger compression, the fine stage revisits the shortened cache at a later denoising step using updated, sampled queries. It ranks feature channels by their mean absolute query activations and uses the highest-scoring half for importance estimation, limiting the additional refinement cost. Retained keys and values remain full-dimensional. Experiments on image editing demonstrate favorable quality–efficiency trade-offs across compression levels, with substantial KV cache reductions and competitive editing quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.