acceptodds
Under review as a conference paper at ICLR 2027

HTSC-AT: HIERARCHICAL TEMPORAL SEMANTIC COMPRESSION WITH ADAPTIVE TOKENIZATION FOR EFFICIENT LONG VIDEO UNDERSTANDING

Abstract

Long video understanding with multimodal large language models (MLLMs) is bottlenecked by quadratic attention complexity and severe temporal redundancy in visual tokens. We present HTSC-AT, a three-stage framework for efficient long-context reasoning on lightweight open-source MLLMs. Stage 1 introduces adaptive visual tokenization driven by motion-semantic change intensity, reducing token count by 66.3% while improving QA accuracy by +5.3%. Stage 2 performs event-aware temporal compression via dual-branch boundary detection and hierarchical intra-event clustering, achieving up to 720× compression while preserving semantically critical evidence. Stage 3 employs a sparse-dense hybrid decoder that maintains dense intra-event attention and structured sparse interevent attention, reducing peak GPU memory from 28.7 GB to 10.4 GB. Across MLVU, VideoMME, LongVideoBench, and LVBench, HTSC-AT achieves stateof-the-art results among open-source small-scale models—outperforming VideoXL2-8B by 4.1 points on MLVU Test—while supporting over 10,000 frames on a single-node 8×H800 GPU setup.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.