UniCache: Task- and Type-Aware KV Cache Compression for Unified Multimodal Models
Abstract
Unified multimodal models combine understanding, generation, and editing within a single network, offering a promising foundation for versatile multimodal applications. However, growing multimodal contexts make KV cache storage and access increasingly costly. Existing KV cache compression methods can reduce these costs, but applying a single policy uniformly can degrade quality across tasks, as shared attention brings together heterogeneous cache types with different attention distributions and temporal dynamics. Based on these findings, we propose UniCache, a training-free framework for task- and type-aware KV cache compression. UniCache identifies the cache segments activated by each task and assigns suitable compression policies to different segment types. It coordinates their parallel execution under a shared storage budget through attention-guided allocation and task-aware temporal scheduling. Experiments show that UniCache maintains balanced quality at high compression ratios: compresses KV cache 5 for understanding and editing and 2.5 for generation with negligible loss, and increases throughput up to 1.78, significantly improving the practicality of scaling unfied multimodal model to longer context length.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.