Adaptive Visual KV Cache Reuse for Efficient Multi-Turn Iterative Multimodal Reasoning
Abstract
Multi-turn visual reasoning acquires evidence through repeated cropping and local inspection, but overlapping observations lead to redundant visual encoding and prefill. We propose an adaptive visual KV cache reuse framework that reduces repeated processing while retaining access to additional image detail. The central difficulty is that skipping visual prefill also removes the interaction between a new observation and the current reasoning: historical keys remain bound to old positions, while historical values have not incorporated the intent formed between crops. Our framework first selects full re-encoding or regional reuse according to crop size and cache availability. On the reuse path, Positional Key Rebinding (\PKR) corrects attention matching at the new positions, and Intent-guided Value Enhancement (\IVE) uses pre-crop reasoning attention to strengthen content relevant to the current objective. Source states from fully re-encoded crops are retained for subsequent retrieval and reuse. The framework requires no additional training; on reuse hits, it bypasses the visual encoder and prefill for the selected visual tokens, then resumes normal text reasoning. Experiments show that our method reduces inference time while preserving the overall accuracy of multi-turn visual reasoning that relies on intermediate crop images.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.