ContextTom: Training-Free Video VLM Acceleration with Early Spatial Context
Abstract
Video vision-language models process thousands of visual tokens before producing an answer. Reducing these tokens early saves computation, but also determines which evidence reaches the language model. We introduce ContextTom, a training-free framework that separates the timing of temporal and spatial compression. Its central idea is to compress time before vision encoding while preserving spatial coverage through early multimodal processing. A fixed-budget temporal partition produces representative medoid anchors, supplemented with bounded local detail from neighboring observations. The anchors retain their complete spatial grids through a short decoder prefix, allowing question tokens to access the visual evidence and completing Qwen3-VL\textquotesingle s multi-level visual injections. ContextTom then selects spatial representatives using query attention and agreement across visual feature levels, and continues directly from their computed states. This schedule reduces vision computation and later decoder workload while reusing the prefix. On Qwen3-VL-8B, ContextTom achieves 76.22% accuracy on Video-MME short with a mean time to first token (TTFT) of 123.5 ms, yielding a speedup over full-input inference. It matches the full-input accuracy of 72.20% on EgoSchema with a speedup and achieves a speedup on LongVideoBench at 59.24% accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.