Saevia: Spatiotemporal Token Compression with Attention Compensation for Efficient Video Understanding
Abstract
Video large language models (VideoLLMs) have demonstrated strong capabilities in video understanding, yet processing numerous visual tokens incurs high inference latency and limits practical deployment. Existing token compression methods reduce temporal and spatial redundancy but often overlook compression-induced attention distortion, which can impair accuracy. We introduce Saevia, a training-free framework that combines spatiotemporal token compression with attention compensation. Specifically, position-wise temporal merging reduces temporal redundancy at each spatial position. Request-aware spatial pruning selects tokens by jointly considering visual salience and request relevance. During LLM prefill, sparsity-aware attention compensation mitigates the attention distortion caused by compression. Across three VideoLLMs and five video understanding benchmarks, Saevia preserves 95.3–99.2% of full-token average accuracy at 10% visual token retention. With dedicated Triton kernels and a FlashAttention-compatible attention compensation design, Saevia achieves 2.6–3.2× time-to-first-token (TTFT) speedups over full-token baselines on one NVIDIA A100 80GB GPU and up to 1.25× end-to-end throughput with 1K-token outputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.