PreViC: Predictive Visual Compression for Video Vision-Language Models
Abstract
Video vision-language models (VLMs) turn even a short video into thousands of visual tokens, and those tokens dominate the prefill that precedes every answer. Cutting them before prefill, without retraining the model, is where compression pays off in deployment. Training-free methods for this stage mostly rank a token by how it looks on its own, using saliency, diversity, or similarity to neighboring frames. Such a score says nothing about how much a token adds to what has already been kept, so it keeps detail that nearby frames already explain and discards small changes that they do not. We propose PreViC (Predictive Visual Compression), which keeps a visual token exactly when a compact visual memory fails to predict it. PreViC is training-free and prompt-agnostic, and runs before the language model. A few temporally novel anchor frames form that memory, and every other token is ranked by how poorly its nearest anchors predict it. Over three checkpoints and five retain ratios, PreViC has the highest four-benchmark average accuracy of any method at 12 of the 15 settings. On LLaVA-OneVision 7B it holds the uncompressed model’s accuracy on 50% of the visual tokens, an operating point no baseline we measured reaches on either 7B model. Large video VLMs can therefore be served far more cheaply without giving up accuracy. More broadly, token compression is better posed not as ranking tokens by how important they look, but as asking what a retained visual memory cannot already predict.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.