acceptodds
Under review as a conference paper at ICLR 2027

ReVisKV: Converting Dynamic Visual Execution Sparsity into Reversible GPU KV Residency

Abstract

Vision-language models repeatedly access visual evidence during extended reasoning, making the visual key-value (KV) cache a persistent source of GPU memory pressure. Dynamic visual selection reduces attention computation, but retaining complete candidate KV on GPU leaves this storage cost unchanged. We introduce ReVisKV, which preserves complete visual KV in CPU memory for managed decoder layers and maintains only a small hot buffer on GPU, recalling missing entries on demand. Motivated by cross-layer drift in visual working sets, ReVisKV refreshes selections at multiple anchors, allowing offloading to begin early while preserving access to later-needed evidence. The hot buffer exploits short-term reuse across decoding steps to reduce transfers, and stage-shared cache decisions with asynchronous execution reduce management overhead. On Qwen3-VL-Thinking 4B, ReVisKV retains 99.5% of VisionPulse's average score across five benchmarks at a 10% selection upper bound, while using 62.5% less GPU visual-KV capacity. Compared with dense decoding, the optimized runtime delivers up to decode speedup and supports up to the batch size.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.