acceptodds
Under review as a conference paper at ICLR 2027

Codec-derived Information Priors for Video Token Sparsification in Vision Language Model Inference

Abstract

We propose a codec-guided video token sparsification framework for high-throughput large vision-language model serving. Existing methods either require full visual encoding for frame selection or depend on intermediate features and attentions for token pruning, placing additional model-side computation on the GPU critical path and increasing serving overhead under concurrent requests. To this end, we leverage compressed-domain information exposed by video codecs to jointly perform frame selection and token mask generation before visual encoding. Specifically, frame selection is guided by coding cost and motion magnitude extracted from the bitstream, while token masks are generated from spatial priors based on motion and texture complexity. By resolving visual input selection on the CPU before visual encoding, IPB decouples sparsification from downstream GPU inference and enables CPU–GPU overlap across requests. At high concurrency, IPB achieves 3.2 the aggregate throughput and 38.0 lower TTFT than a combined frame-selection and token-pruning baseline, while maintaining favorable accuracy–efficiency trade-offs across four video benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.