acceptodds
Under review as a conference paper at ICLR 2027

Visual Tokens Should Not Be Pruned Forever: Budgeted Visual Token Refresh for Multi-Turn LVLMs

Abstract

Large vision-language models (LVLMs) often retain only a subset of visual tokens to reduce inference costs. However, existing static pruning methods typically select tokens according to the initial question and reuse the same token set throughout subsequent dialogue turns. Consequently, when the focus of a conversation changes, visual information required by a new question may remain inaccessible even though it is present in the image. We formulate this limitation as the problem of adapting the active visual context to follow-up questions under a fixed token budget and propose Budgeted Visual Token Refresh (BVTR). BVTR initially activates 288 of 576 visual tokens and, upon receiving a follow-up question, replaces up to 48 tokens according to their question-conditioned importance. Because directly inserting KV states can mix representations computed under inconsistent visual contexts, BVTR-Dense reconstructs the preceding context using the refreshed token set and replays the answer tokens generated in the previous turn. BVTR-Sparse further reduces reconstruction costs by preserving the original token order and selectively updating only a subset of deep-layer KV states. On the frozen COCO-TS1000 evaluation with LLaVA-1.5-7B, BVTR-Dense improves accuracy over the fixed 288-token baseline by 1.4 percentage points, with a significant paired difference under McNemar’s test (). In a four-turn topic-switching evaluation, it increases the proportion of conversations in which all three follow-up questions are answered correctly from 68.0% to 75.5% (), while also yielding a significant improvement on an MME multi-turn adaptation. On a separate confirmatory evaluation of 1,000 images, BVTR-Sparse improves accuracy over the fixed baseline by 2.1 percentage points (). It reduces incremental peak GPU memory by 35.4% while achieving an observed accuracy only 0.1 percentage points below that of the full 576-token context, and reduces paired end-to-end latency by 15.9% relative to BVTR-Dense. These results demonstrate that adapting the visual context to shifts in conversational focus under a fixed token budget provides an effective trade-off between accuracy and efficiency for multi-turn LVLMs, distinguishing BVTR from static pruning approaches that treat visual-token removal as a permanent decision.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.