Who Does What, and When? Synchronized Conjunctive Reasoning for Image-Disambiguated Video Temporal Grounding
Abstract
Image-Disambiguated Video Temporal Grounding (ID-VTG) aims to localize a temporal segment using a reference image that identifies the target entity and a text query that describes the target event. The task is inherently conjunctive: entity and event conditions must be supported at the same temporal location, while satisfying either condition alone is insufficient. Existing ID-VTG methods effectively integrate image and text cues, yet their entangled proposal relevance leaves temporal entity–event conjunction implicit, limiting their ability to reliably determine when both conditions hold simultaneously. To address this problem, we first introduce Synchronized Conjunctive Grounding (SCG), which factorizes frame-wise entity and event evidence and models their temporal conjunction through a non-compensatory product. The resulting conjunctive evidence is propagated across the temporal pyramid and injected as a signed residual to calibrate proposal relevance, favoring moments jointly supported by both conditions. With entity–event conjunction explicitly modeled, we further characterize its underlying validity structure through Factor-Selective Counterfactual Learning (FCL), making the validity of entity and event factors identifiable. Rather than collapsing intervened queries into generic negatives, FCL exposes the underlying factorial validity structure of entity and event conditions and learns a structured four-state energy with unary validity and pairwise interaction potentials. Together, SCG and FCL form a coherent factorize–conjoin–attribute framework for effective entity–event conjunctive grounding. Extensive experiments on existing ID-VTG benchmarks validate the effectiveness of the proposed framework, with consistent improvements across temporal granularities, ambiguity settings, and cross-domain generalization scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.