HypToC: Hyperbolic Vision Token Compression for Video Large Language Models
Abstract
Video large language models (Video-LLMs) achieve strong video understanding capabilities, while the large number of vision tokens substantially increases inference cost. Recent training-free compression methods have attained promising efficiency by modeling token importance and redundancy in video representations. In this study, we investigate the geometry of video token representations and find that their embedding space exhibits pronounced tree-like organization compatible with hyperbolic geometry. Within individual frames, deeper vision features reveal increasing spatial hierarchy, whereas between frames, tokens corresponding to the same semantic object maintain a coherent temporal hierarchy despite substantial spatial displacement. Motivated by these findings, we introduce HypToC, a training-free framework that exploits hyperbolic geometry for video token compression with two new modules. Intra-Frame Lorentz Selection retains a fixed-size set of informative and geometrically diverse tokens. Inter-Frame Lorentz Clustering then aggregates the remaining tokens according to continuous lowest common ancestor relations in hyperbolic space. HypToC is training-free, plug-and-play, and model-agnostic. Experiments across representative Video-LLMs and benchmarks show that HypToC achieves state-of-the-art average performance under diverse matched token budgets.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.