STORS: Spatio-Temporal Orthogonal Residual Selection for Video Token Compression
Abstract
Efficient video understanding requires compressing highly redundant visual tokens across space and time. Existing geometric selection methods typically operate on decoder-facing representations, overlooking how the visual projector may alter the geometry on which their selection criteria depend. We identify a substantial change in representation geometry across the projector in Qwen2.5-VL: encoder-side features exhibit extreme anisotropy and near-saturated pairwise similarities, whereas decoder-facing features are substantially more dispersed. Motivated by this finding, we propose STORS, a training-free method that scores tokens by their unexplained residual energy over a unified spatio-temporal pool while retaining their projected features for the language model. An explicit temporal redundancy weight controls cross-frame redundancy under a fixed token budget. Controlled experiments reveal that the effectiveness of residual selection depends strongly on the scoring space: on Video-MME, the same criterion underperforms uniform grid sampling in all nine settings when applied after the projector, but outperforms it in eight of nine when applied before the projector. In contrast, no statistically detectable scoring-space benefit is observed in LLaVA-OneVision, whose representations exhibit similar geometry. Across five video benchmarks and two backbones, STORS achieves competitive performance under aggressive compression. On Qwen2.5-VL-7B, retaining 10% of visual tokens preserves 94.4% of the uncompressed average accuracy while achieving a speedup in time to first token.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.