acceptodds
Under review as a conference paper at ICLR 2027

DART: Decoupled Axis-Adaptive Reconstruction Trees for Budgeted Video Token Compression

Abstract

Video large language models incur substantial computational costs when process- ing dense frame-wise visual tokens. Compressing these tokens under a fixed bud- get requires deciding not only which regions to preserve, but also whether to retain temporal variation, spatial detail, or both. We formulate this decision as budgeted, attention-weighted reconstruction of the video feature volume and propose Decou- pled Axis-Adaptive Reconstruction Trees (DART). DART represents the video as a partition of contiguous spatiotemporal regions, with one pooled token per re- gion. A greedy planner applies temporal, spatial, and spatiotemporal partition- ing to every candidate region, evaluates the reconstruction error of each resulting partition with respect to the original features within the region, and selects the re- gion and partitioning strategy that yield the greatest reconstruction-error reduction per additional token. Adaptive-T further selects content-dependent boundaries for temporal-only splits, while spatial and spatiotemporal splits use midpoint bound- aries. The resulting partition guides importance-weighted pooling of the original features before language-model prefilling, without query conditioning or updates to pretrained parameters. Averaged over four video-understanding benchmarks, DART retains 101.2% and 102.7% of the corresponding full-token baseline ac- curacy on LLaVA-OneVision-7B and 0.5B, respectively, while using only 25% of the visual tokens. These results support budgeted region refinement as an effective approach to allocating spatiotemporal resolution for video token compression.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.