FoveaTree: Information-Density Trees for Query-Guided Visual Token Pruning in Video Large Language Models
Abstract
Video large language models (Video LLMs) have demonstrated strong multimodal understanding across a range of tasks, including visual question answering and video captioning. Yet, encoding each frame into hundreds of visual tokens makes inference over temporal visual content expensive. Existing methods reduce this cost by merging similar tokens or using text-to-vision attention to select task-relevant tokens. Nevertheless, these methods still face two main challenges: (1) Hierarchy-blind Pruning: many similarity-based methods rely on pairwise similarity or training a lightweight neural network to predict importance scores without explicitly considering the regional compressibility of a frame, which can misguide the merging or removal of critical evidence. (2) Attention Drift: position-dependent effects from Rotary Position Embedding (RoPE), which encodes token positions in the decoder, together with preferences for nearby or recent tokens and attention sinks can shift query-to-vision attention away from task relevance, causing attention-based pruning to underweight and discard task-relevant visual tokens. Motivated by these observations, we propose FoveaTree, a two-stage, training-free visual-token pruning framework for efficient video inference. The first stage estimates regional compressibility from the feature distortion incurred when merging a region to construct a quadtree-derived density map, which guides the dropping, merging, and retention of visual regions. Based on the surviving units and inherited tree topology, the second stage introduces an efficient tree-structured pruning mechanism in the middle layers of the LLM to mitigate Attention Drift and further reduce low-relevance visual content using text-to-vision attention. We validate FoveaTree under different Video LLM architectures, i.e., InternVL3.5-8B and Qwen3-VL-8B, on five video question-answering benchmarks and one captioning benchmark with token budgets of 50%, 30%, and 10%; our method consistently achieves overall state-of-the-art performance across all evaluated settings, demonstrating a superior accuracy-efficiency trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.