HTSC-AT: HIERARCHICAL TEMPORAL SEMANTIC COMPRESSION WITH ADAPTIVE TOKENIZATION FOR EFFICIENT LONG VIDEO UNDERSTANDING
Abstract
Long video understanding with multimodal large language models (MLLMs) is bottlenecked by quadratic attention complexity and severe temporal redundancy in visual tokens. We present HTSC-AT, a three-stage framework for efficient long-context reasoning on lightweight open-source MLLMs. Stage 1 introduces adaptive visual tokenization driven by motion-semantic change intensity, reducing token count by 66.3% while improving QA accuracy by +5.3%. Stage 2 performs event-aware temporal compression via dual-branch boundary detection and hierarchical intra-event clustering, achieving up to 720× compression while preserving semantically critical evidence. Stage 3 employs a sparse-dense hybrid decoder that maintains dense intra-event attention and structured sparse interevent attention, reducing peak GPU memory from 28.7 GB to 10.4 GB. Across MLVU, VideoMME, LongVideoBench, and LVBench, HTSC-AT achieves stateof-the-art results among open-source small-scale models—outperforming VideoXL2-8B by 4.1 points on MLVU Test—while supporting over 10,000 frames on a single-node 8×H800 GPU setup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.