STEER: State-Transition Event Evidence Reasoning for Training-Free Video Temporal Grounding
Abstract
Video Temporal Grounding (VTG) aims to localize the temporal boundaries of a target event in an untrimmed video according to a natural language query. Exist- ing training-free methods typically identify target moments based on video-query semantic relevance. However, semantic relevance does not necessarily indicate that the target event is actually occurring. Moments before and after the event may contain similar persons, objects, and scenes, leading to ambiguous temporal boundaries. To address this issue, we propose STEER (State-Transition Event Evidence Reasoning), a training-free framework that explicitly reasons event pro- gression through complementary semantic, motion, and query-independent cap- tion evidence. STEER first parses the query into a structured event attributes, which guides the construction of complementary evidence from video-text seman- tics, action-aware human motion, and query-independent video captions. These cues are jointly used for event-state reasoning, and the resulting state transitions are used to predict the temporal boundaries. Experiments on Charades-STA and ActivityNet Captions demonstrate competitive performance, while ablation stud- ies further validate the effectiveness of the proposed evidence construction and state-transition reasoning. Code and related resources will be publicly available upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.