DART: Decoupled Axis-Adaptive Reconstruction Trees for Budgeted Video Token Compression
Abstract
Video large language models incur substantial computational costs when process- ing dense frame-wise visual tokens. Compressing these tokens under a fixed bud- get requires deciding not only which regions to preserve, but also whether to retain temporal variation, spatial detail, or both. We formulate this decision as budgeted, attention-weighted reconstruction of the video feature volume and propose Decou- pled Axis-Adaptive Reconstruction Trees (DART). DART represents the video as a partition of contiguous spatiotemporal regions, with one pooled token per re- gion. A greedy planner applies temporal, spatial, and spatiotemporal partition- ing to every candidate region, evaluates the reconstruction error of each resulting partition with respect to the original features within the region, and selects the re- gion and partitioning strategy that yield the greatest reconstruction-error reduction per additional token. Adaptive-T further selects content-dependent boundaries for temporal-only splits, while spatial and spatiotemporal splits use midpoint bound- aries. The resulting partition guides importance-weighted pooling of the original features before language-model prefilling, without query conditioning or updates to pretrained parameters. Averaged over four video-understanding benchmarks, DART retains 101.2% and 102.7% of the corresponding full-token baseline ac- curacy on LLaVA-OneVision-7B and 0.5B, respectively, while using only 25% of the visual tokens. These results support budgeted region refinement as an effective approach to allocating spatiotemporal resolution for video token compression.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.