TSL:Learning the Temporal Structure of Evidence via State Flow for Temporal Video Understanding
Abstract
Video temporal understanding integrates visual and linguistic information, requiring models to answer accurately while capturing when events occur and how they relate over time. Semantically related evidence segments form complete evidence chains with temporal structure, yet existing methods that directly predict evidence interval sets inadequately model temporal relations and gaps between segments, leading to missed evidence segments, erroneous splitting, and boundary localization errors. To address these challenges, we propose a method for learning temporal evidence structure through state flow, explicitly modeling the onset, continuation, and termination of query-relevant evidence along the timeline as a learnable evidence-state transition process that organizes local evidence segments into structurally complete chains. Our theoretical analysis establishes a reversible correspondence between evidence interval sets and canonical evidence state–offset representations. To learn this representation, we design a Temporal Structure Learner (TSL) that injects information about temporal evidence structure into video representations between the decoder layers of a multimodal large language model. Through supervised learning, the model predicts legal evidence-state chains and refines the endpoints of predicted segments using offsets, enabling discrete state chains and continuous boundary offsets to jointly determine complete evidence chains. Experimental results demonstrate that our method substantially outperforms existing approaches to video temporal understanding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.