Rotation-Aligned Key Channel Pruning for Efficient Vision-Language Model Inference
Abstract
Vision-Language Models suffer severe KV cache pressure at inference, as a single image often encodes into thousands of tokens. To reduce this cost, most prior works exploit token sparsity through token pruning, but permanently discarding visual content substantially degrades fine-grained perception. This motivates a complementary axis, feature sparsity, which compresses the channel dimension to preserve more visual tokens at the same KV cache budget. Existing Key channel pruning methods, however, face a structural trade-off: token-wise channel pruning is accurate but unstructured and slow, while head-wise pruning is hardware-friendly but degrades at high compression. We attribute this degradation to informative Key directions that change across tokens and inputs, so no fixed channel subset serves them all. We propose RotateK, a training-free method that builds a query-weighted Key basis for each input at prefill, allowing a head-wise mask to retain attention-relevant information. To make this practical, we implement a graph-capturable top-k eigensolver, a fused sparse-channel decode kernel, and a paged KV layout in vLLM. Across six VLMs, up to 32B parameters, and four token pruning methods, RotateK consistently improves over token pruning alone at matched KV budgets and is comparable in accuracy to token-wise channel pruning, while increasing KV capacity and generation throughput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.