acceptodds
Under review as a conference paper at ICLR 2027

ViCoS: Verify-then-Correct Retrieval and Context Mosaic for Aggressive Video Token Compression

Abstract

Long-form video question answering requires vision-language models (VLMs) to process thousands of visual tokens, making inference expensive due to the quadratic cost of self-attention. Training-free visual token compression reduces this cost, but performance can degrade substantially under aggressive budgets. We identify two related problems: mis-selection, where the limited visual budget is spent on answer-irrelevant frames, and discarded-context, where unselected frames retain substantial query-aligned visual signal but are removed entirely. We propose ViCoS (Verify-then-Correct with Context Mosaic), a training-free framework that addresses both problems while keeping the final answer-stage visual-token budget fixed.VerCoRe self-verifies the current retrieved frames and, when the evidence is insufficient, generates corrective visual concepts to refine CLIP retrieval. CoMo preserves discarded temporal context by packing the remaining candidate frames into a single low-resolution mosaic. Under a budget of K images, ViCoS uses K-1 full-resolution focus frames and one mosaic; at K=1, it reduces to CoMo alone. Across two backbones and five video QA benchmarks, ViCoS achieves the highest compressed accuracy among the compared methods in all 30 evaluated backbone–benchmark–budget settings. On VideoVista with InternVL3-8B, ViCoS achieves 78.80 accuracy at a 12.5% visual-token budget, only 0.04 points below the full K=8 reference while using a single encoder image.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.