Visual Evidence-based Reasoning for Multi-Object Multi-State Transitions
Abstract
Understanding object states in videos requires not only recognizing physical properties but also linking state changes to specific actions and conditions. Recognizing states and actions does not ensure correct object–state associations or causal explanations: Even when models correctly identify states, the explanations they provide may not align with the actual video content despite potentially conforming to common sense. Models therefore need to reason from visual evidence to connect each state change to the specific operations and conditions affecting the corresponding object. To address this, we propose a graph-guided visual evidence-based reasoning framework for understanding multi-object multi-state transitions in videos. A state graph is built based on timestamped segment descriptions and visual observations through iterative relation verification and state completion by leveraging a Vision-Language Model (VLM). In the graph, nodes represent objects and their states at specific stages, and edges represent object state transitions, multi-object combinations, or single-object separations. This structure links each state change to the corresponding object, relevant actions, and visual observations supporting these associations. For each transition, the VLM infers potential causes from object and contextual information in connected nodes and edges, accounting for implicit actions whose effects are difficult to observe locally. The timestamps retained in the graph are used to locate visual evidence and verify the consistency between state differences across nodes and actions represented by edges, thereby correcting erroneous object–state associations and revising or rejecting unsupported causal explanations. To support this reasoning framework, we construct VERMOST, a multi-object multi-state dataset, which comprises 1,827 third-person cooking videos and 75,977 annotated state transition events. Annotations include objects, pre- and post-transition states, a temporal interval, relevant actions, and specific causes along with their types. VERMOST links states and transitions to video-supported causes, covering explicit actions, implicit actions, and their joint effects. Experiments on VERMOST show that our framework outperforms the baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.