DynaSense: Learning to Perceive Dynamic Changes in Videos for Temporal Grounding
Abstract
Current video Multi-modal Large Language Models (Video-MLLMs) for Video Temporal Grounding (VTG) rely on RGB-centric static visual representations, which capture visual content but offer limited modeling of visual state evolution over time. Consequently, these approaches fail to perceive dynamic visual changes at the temporal boundaries of events, where fine-grained temporal transitions are critical for accurate localization. To address this limitation, we propose DynaSense, a Video-MLLM that improves temporal grounding by explicitly modeling dynamic representations of temporal changes within videos. DynaSense models continuously evolving videos by representing visual state changes at each temporal step relative to the preceding step. It first aligns visual features at the current timestamp with the most relevant references from the preceding timestamp within a position-free embedding space. Subsequently, it encodes the resulting temporal change as the difference between the corresponding visual features after incorporating 2D spatial positional encodings. This dynamic representation captures both appearance differences and spatial displacements, providing temporally discriminative evidence for localizing event boundaries. Furthermore, we introduce video state change recognition dataset (VideoSCR), a video state-change recognition dataset constructed via an automated annotation pipeline, to enhance the perception and recognition of video dynamics. DynaSense consists of two training stages. The first performs state-change question answering on VideoSCR to facilitate the learning of dynamic video features, and the second introduces video temporal grounding data to strengthen the capability for temporal localization. Experiments across multiple VTG benchmarks demonstrate that DynaSense consistently improves temporal grounding accuracy and can be readily integrated into other Video-MLLMs to enhance their temporal grounding capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.