acceptodds
Under review as a conference paper at ICLR 2027

A Dynamic Temporal Segmentation for Long Video Understanding

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable success in video understanding, yet scaling them to long videos remains a major challenge due to severe temporal redundancy and limited frame budgets. Some existing methods rely on uniform temporal segmentation, implicitly assuming that semantic content varies uniformly over time, which misallocates frame budgets and misses critical moments. In this paper, we propose DyQCA, a Dynamic Segmentation framework for keyframe selection in long video understanding. DyQCA detects temporal boundaries by comparing the semantic differences of adjacent frames, partitions the video into dynamic segments, and allocates the frame budget across segments by jointly modeling query relevance and content deviation, followed by anchor-centric keyframe selection within each segment. The proposed DyQCA is training-free, model-agnostic, and can be seamlessly plugged into existing Video-LLMs. Extensive experiments demonstrate that DyQCA consistently outperforms uniform sampling and recent keyframe selection baselines, e.g., 67.3% on LongVideoBench and 76.2% on MLVU under a 64-frame budget, demonstrating that dynamic semantic segmentation offers a simple yet effective path toward more accurate long video understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.